Cloud Operations Suite: Logging, Sinks, Monitoring & Alerts

Master GCP Cloud Logging, Audit vs Data Access logs, building Log Router Sinks to BigQuery & GCS, log-based metrics, and Cloud Monitoring alert policies.

advanced 22 min lesson hands-on task included

GCP’s Cloud Operations Suite (formerly Stackdriver) provides real-time logging, monitoring, APM tracing, and error reporting across GCP, AWS, and on-premises environments.

In high-scale production systems, unmanaged logging is one of the single largest sources of cloud bill bloat. Understanding Log Router Sinks and Exclusion Filters is critical for both security and FinOps.


Topic 1: Cloud Logging Architecture & Log Categories

CLOUD LOGGING ROUTER & LOG SINKS PIPELINE LOG SOURCES • App Stdout / Stderr • GCE Ops Agent • Audit Logs (Admin Activity) • Data Access Logs • VPC Flow Logs • GKE FluentBit / Vector LOG ROUTER & FILTERS resource.type="gce_instance" severity>=ERROR Evaluates log entries in real-time against defined Log Sink query filters. • _Default Sink (30-day) • Exclusions (Cost control) SINK DESTINATIONS Cloud Storage (Long Archive) BigQuery (SQL Analytics) Pub/Sub (SIEM / Datadog) LOG-BASED METRICS & CLOUD MONITORING ALERTS Log-based Counter / Distribution Metric Converts regex log occurrences into a time-series metric Alerting Policy & Notification Channels Triggers PagerDuty, Slack, Email on error threshold breach FINOPS TIP: Log ingestion is billed per GB. Use Exclusion Filters to drop noisy debug logs before ingestion, saving thousands per month.
GCP Log Router Sink pipeline: Ingesting app/audit logs, evaluating Log Router query filters, exporting to BigQuery/GCS/PubSub, and triggering Cloud Monitoring alerts.

GCP categorizes logs into two primary categories:

  1. User Logs:
    • Generated by applications, services, and workloads. Written via the Cloud Logging API, Ops Agent, or GKE FluentBit/Vector daemonsets.
  2. Security & System Logs:
    • Admin Activity Audit Logs: Records administrative configuration changes (e.g., creating a VM, modifying IAM rules). Enabled by default, free of charge, retained for 400 days.
    • Data Access Audit Logs: Records API calls that inspect or modify user data (e.g., reading a GCS object, querying a BigQuery table). Must be explicitly enabled; incurs log storage costs.
    • Access Transparency Logs: Records actions taken by Google staff when responding to support tickets.

Topic 2: Log Router Sinks & Real-Time Export Pipelines

By default, Cloud Logging stores logs in the _Default log bucket for 30 days. To retain logs for compliance or analyze them with SQL, you build Log Router Sinks:

A Log Sink routes matching log entries to one of three primary destinations:

  • Google Cloud Storage (GCS): Bulk long-term archival (JSON files batch-written hourly). Lowest storage cost.
  • BigQuery: Real-time streaming into columnar SQL tables for instant querying and security analytics.
  • Pub/Sub: Real-time event streaming to third-party SIEM or observability platforms (Datadog, Splunk, Elastic).
# Create a Log Router Sink exporting all ERROR logs to BigQuery
gcloud logging sinks create sink-prod-errors-bq \
  bigquery.googleapis.com/projects/my-project-id/datasets/sec_log_archive \
  --log-filter='severity>=ERROR AND resource.type="gke_container"'

# FinOps Best Practice: Create an Exclusion Filter to drop noisy debug logs BEFORE ingestion
gcloud logging sinks update _Default \
  --add-exclusion="name=drop-debug-logs,filter=severity=DEBUG"

Topic 3: Log-Based Metrics & Cloud Monitoring Alerts

When an application produces specific text patterns in stdout (e.g., FATAL: Connection pool exhausted), you can convert those log occurrences into time-series metrics using Log-Based Metrics:

  1. Counter Metrics: Counts the number of log entries matching a filter over time.
  2. Distribution Metrics: Extracts numerical values from log payloads (e.g., latency values) to calculate latencies, P99s, or payload sizes.
# Create a Log-Based Counter Metric tracking database connection failures
gcloud logging metrics create db_conn_failures \
  --description="Counts database connection pool failures" \
  --log-filter="textPayload:\"Connection pool exhausted\""

Topic 4: Cloud Monitoring Workspace & Alert Policies

Cloud Monitoring aggregates metrics from GCE, GKE, Cloud SQL, and custom log-based metrics:

  • Alerting Policies: Define Conditions (e.g., db_conn_failures > 5 for 5 minutes), Duration, and Notification Channels (Slack, PagerDuty, Email, Pub/Sub webhook).
  • Cloud Trace: Distributed APM tracing that measures latency across microservice HTTP calls (automatically integrated into App Engine and GKE).
  • Error Reporting: Aggregates stack traces from Go, Python, Java, Node.js, and displays grouped crash frequency dashboards.

Common mistake: Leaving the _Default bucket’s 30-day retention in place and calling it a logging strategy. Audit logs you will need for an investigation expire, high-volume application logs you will never read are retained at full price, and neither decision was made. Route with sinks — audit logs to a locked bucket in a separate project, noise to Cloud Storage or nowhere.