Observability: CloudWatch, CloudTrail and Config

Three services that answer three different questions, the metric-filter and composite-alarm patterns worth building once, and why alarming on utilisation produces pages nobody acts on.

advanced 22 min lesson hands-on task included

AWS gives you three observability services with overlapping names and non-overlapping jobs. Using the wrong one wastes hours during an incident, and using all three without knowing what each is for produces a monitoring bill that rivals the compute it watches.


Topic 1: Three Services, Three Questions

THREE SERVICES, THREE DIFFERENT QUESTIONS — DO NOT SUBSTITUTE ONE FOR ANOTHER CloudWatch "Is it healthy right now?" metrics · 1s–5min granularity logs + metric filters alarms → SNS / Auto Scaling dashboards, Logs Insights BLIND SPOT Tells you nothing about who changed what. CloudTrail "Who did this, and when?" every management API call 90-day event history, free a trail → S3 for longer data events are opt-in, per resource BLIND SPOT Not real-time: events land in ~5–15 min. AWS Config "What did it look like then?" resource configuration snapshots timeline of every change rules → compliant / not conformance packs BLIND SPOT Per-item pricing; recording everything gets expensive. THE SEQUENCE IN AN ACTUAL INCIDENT CloudWatch alarm fires → Logs Insights narrows the window → CloudTrail names the principal and the call → Config shows the before-and-after. Skipping straight to CloudTrail on a 40-million-event day is how a 20-minute investigation becomes a 3-hour one.
Read the blind spots at the bottom of each column. Each service is authoritative for exactly one question and useless for the other two, which is why the incident sequence runs left to right rather than starting wherever the console opens.

CloudWatch — “is it healthy right now?” Metrics, logs, alarms, dashboards. Everything numeric, at 1-minute standard resolution or 1-second for high-resolution custom metrics.

CloudTrail — “who did this?” Every management API call in the account: the principal, the source IP, the parameters, the response. The last 90 days are free in Event History; longer retention needs a trail delivering to S3.

AWS Config — “what did it look like then?” Point-in-time configuration snapshots of resources, a change timeline, and rules that mark resources compliant or not.

The sequence during an incident is the reason to know the difference:

1. CloudWatch alarm fires            → something is wrong, and when it started
2. Logs Insights on that window      → what the application was saying
3. CloudTrail for that window        → who or what changed something
4. Config timeline for that resource → exactly what changed, before and after

Starting at step 3 on a busy account means paging through tens of millions of events. Narrowing with steps 1 and 2 first turns that into a targeted lookup.


Topic 2: Metrics That Are Worth an Alarm

Alarm on symptoms, not on utilisation. High CPU is not an incident; users getting errors is an incident. A CPU alarm pages you for a batch job doing its job, and stays silent while a deadlocked process serves 500s at 4% CPU.

The symptom metrics that map to user pain:

LayerMetricWhy
ALBTargetResponseTime p99Latency as users experience it
ALBHTTPCode_Target_5XX_CountErrors your application produced
ALBHTTPCode_ELB_5XX_CountErrors the load balancer produced — usually no healthy targets
ALBHealthyHostCountCapacity about to become an outage
SQSApproximateAgeOfOldestMessageBetter than queue depth: it is time, not count
RDSReadLatency, DatabaseConnectionsThe two that precede most database incidents
LambdaErrors, Throttles, IteratorAgeFailure, capacity, and falling behind
AnyCustom business metricOrders per minute detects what no infrastructure metric can

Percentiles, not averages. An average hides the tail entirely: a service where 1% of requests take 30 seconds has a fine average and a broken experience. CloudWatch supports p50, p90, p99 and arbitrary percentiles on any metric with the statistic set accordingly.

Metric math turns raw counters into meaningful ratios, which is usually what you actually want to alarm on:

# Error RATE, not error count — immune to traffic volume changes
e1 = (m1 / m2) * 100
  m1 = HTTPCode_Target_5XX_Count (Sum)
  m2 = RequestCount (Sum)
alarm when e1 > 1 for 2 datapoints within 5 minutes

An absolute error-count alarm fires during a traffic spike and stays silent during a quiet outage. A rate does neither.

Treat missing data deliberately. TreatMissingData defaults to missing, which means a metric that stops being published leaves the alarm in its last state — an alarm that never fires because the thing it watched disappeared. For a metric that must always exist, breaching is the honest setting.

Composite alarms cut noise by requiring several conditions at once: alarm only when latency is high and the error rate is up and it is not a known deployment window. They also let you suppress downstream alarms while a parent alarm is active, so one incident produces one page rather than forty.


Topic 3: Logs Without an Unbounded Bill

CloudWatch Logs charges for ingestion, storage and query. The ingestion charge is the one that surprises people, because it is per gigabyte and applies to everything including the debug logging someone turned on temporarily.

Four controls, in order of impact:

  1. Set retention on every log group. The default is never expire, which means paying for 2019 logs in 2026. Most application logs are worthless after 30 days; audit logs have a required period; almost nothing needs infinite.
# Find the log groups with no retention set — this list is always longer than expected
aws logs describe-log-groups \
  --query 'logGroups[?!not_null(retentionInDays)].[logGroupName,storedBytes]' --output table
aws logs put-retention-policy --log-group-name /aws/lambda/api --retention-in-days 30
  1. Log at the right level. Debug logging in production is the single largest source of log spend, and it is usually left on after an investigation.
  2. Structure your logs as JSON. Insights parses JSON natively, so structured logs are both cheaper to query and possible to aggregate. Unstructured logs force regex extraction on every query.
  3. Archive cold logs to S3 and query with Athena. Long retention in CloudWatch Logs costs several times S3 for data you touch twice a year.

Metric filters turn a log pattern into a metric you can alarm on — the bridge between logs and alarms:

aws logs put-metric-filter --log-group-name /aws/lambda/api \
  --filter-name payment-failures \
  --filter-pattern '{ $.level = "ERROR" && $.event = "payment_declined" }' \
  --metric-transformations \
    metricName=PaymentFailures,metricNamespace=App,metricValue=1,defaultValue=0

defaultValue=0 matters: without it the metric has no data points when nothing matches, and an alarm on missing data behaves unpredictably.

Logs Insights is the query engine, and knowing three patterns covers most incidents:

fields @timestamp, @message
| filter @message like /ERROR/
| sort @timestamp desc
| limit 50

stats count() by bin(5m)           # when did it start?

filter status >= 500
| stats count() as errors by path
| sort errors desc                  # which endpoint is failing?

Queries are billed by bytes scanned, so always constrain the time range first. That is also the fastest way to get an answer.


Topic 4: CloudTrail, Set Up Properly

Event History is on by default, free, and covers 90 days of management events — enough for most “who changed this” questions and nothing else.

For anything beyond that, create an organization trail in the management account, delivering to an S3 bucket in a separate log archive account:

Trail (organization-wide, all regions)
  → S3 bucket in the log-archive account
      · bucket policy denying delete to everyone
      · Object Lock in compliance mode for the required period
      · SSE-KMS with a key the source accounts cannot use to delete
  → CloudWatch Logs (optional, for metric filters and alarms)

The separate account is the point: an attacker with administrator access in the workload account cannot erase the record of what they did.

Management events vs data events. Management events (CreateBucket, RunInstances, AssumeRole) are on by default. Data events — GetObject on a specific bucket, Invoke on a specific Lambda — are opt-in, per resource, and expensive at volume because they are per-request. Enable them for the buckets that hold sensitive data, not for everything.

Two alarms worth building from CloudTrail metric filters, because they detect the events that precede a bad day: root account usage, and any StopLogging or DeleteTrail call.

CloudTrail Lake is the newer option: an immutable event store you query with SQL, up to a decade of retention, no S3-plus-Athena plumbing. It costs more per event and removes a lot of undifferentiated work.


Topic 5: Config as a Compliance Timeline

Config records resource configuration over time. The two things it gives you that nothing else does:

  • A timeline per resource. “What did this security group look like before Tuesday” answered exactly, with a diff.
  • Rules with automatic remediation. s3-bucket-public-read-prohibited, encrypted-volumes, rds-instance-public-access-check, and your own rules backed by Lambda. Attach an SSM Automation document and a violation is corrected without a human.

Config is priced per configuration item recorded, so recording every resource type in every region in every account gets expensive quickly. Record the resource types that matter — IAM, security groups, S3, EBS, RDS, EC2 — and use conformance packs to deploy a standard set of rules across the organization from a delegated administrator account.


Topic 6: X-Ray, the Agent, and What to Instrument

X-Ray traces a request across services and shows where the time went. On a distributed system it answers the question metrics cannot: not “is the system slow” but “which of these eleven calls is slow, and is it slow for everyone or only for this tenant”. Sampling keeps the cost sane — trace a percentage plus every error.

The CloudWatch agent is what gets you memory and disk metrics from EC2. Those are not available by default, because the hypervisor cannot see inside the instance — which is why “memory utilisation is missing from CloudWatch” is a permanent FAQ. The agent also ships logs, so it is usually the one thing you install on every instance.

Container Insights does the same for EKS and ECS, with pod- and cluster-level metrics. Application Signals builds service-level objectives on top of the traces and metrics — useful when you have moved from “is it up” to “are we meeting the SLO we promised”.

What to instrument first, if you are starting from nothing:

1. ALB metrics and alarms          (free, immediate, symptom-level)
2. Retention on every log group    (saves money on day one)
3. Structured application logs     (makes everything else possible)
4. CloudTrail organization trail   (you cannot backfill an audit trail)
5. A dashboard per service         (the one you open first at 3am)
6. X-Ray on the request path       (once you have more than three services)

Try it yourself: build the error-rate alarm from Topic 2 using metric math, then deploy a deliberately broken version and watch the sequence — alarm, Insights, CloudTrail, Config. Doing it once in calm conditions is what makes it fast in real ones.

Common mistake: creating dozens of alarms on infrastructure utilisation and none on user-facing symptoms. The result is an on-call rotation that is paged constantly for things that are not incidents, learns to ignore the pages, and then misses the one that mattered. Fewer alarms, all of them meaning “a user is affected”, is strictly better than complete coverage nobody trusts.