GCP’s Cloud Operations Suite (formerly Stackdriver) provides real-time logging, monitoring, APM tracing, and error reporting across GCP, AWS, and on-premises environments.
In high-scale production systems, unmanaged logging is one of the single largest sources of cloud bill bloat. Understanding Log Router Sinks and Exclusion Filters is critical for both security and FinOps.
Topic 1: Cloud Logging Architecture & Log Categories
GCP categorizes logs into two primary categories:
- User Logs:
- Generated by applications, services, and workloads. Written via the Cloud Logging API, Ops Agent, or GKE FluentBit/Vector daemonsets.
- Security & System Logs:
- Admin Activity Audit Logs: Records administrative configuration changes (e.g., creating a VM, modifying IAM rules). Enabled by default, free of charge, retained for 400 days.
- Data Access Audit Logs: Records API calls that inspect or modify user data (e.g., reading a GCS object, querying a BigQuery table). Must be explicitly enabled; incurs log storage costs.
- Access Transparency Logs: Records actions taken by Google staff when responding to support tickets.
Topic 2: Log Router Sinks & Real-Time Export Pipelines
By default, Cloud Logging stores logs in the _Default log bucket for 30 days. To retain logs for compliance or analyze them with SQL, you build Log Router Sinks:
A Log Sink routes matching log entries to one of three primary destinations:
- Google Cloud Storage (GCS): Bulk long-term archival (JSON files batch-written hourly). Lowest storage cost.
- BigQuery: Real-time streaming into columnar SQL tables for instant querying and security analytics.
- Pub/Sub: Real-time event streaming to third-party SIEM or observability platforms (Datadog, Splunk, Elastic).
# Create a Log Router Sink exporting all ERROR logs to BigQuery
gcloud logging sinks create sink-prod-errors-bq \
bigquery.googleapis.com/projects/my-project-id/datasets/sec_log_archive \
--log-filter='severity>=ERROR AND resource.type="gke_container"'
# FinOps Best Practice: Create an Exclusion Filter to drop noisy debug logs BEFORE ingestion
gcloud logging sinks update _Default \
--add-exclusion="name=drop-debug-logs,filter=severity=DEBUG"
Topic 3: Log-Based Metrics & Cloud Monitoring Alerts
When an application produces specific text patterns in stdout (e.g., FATAL: Connection pool exhausted), you can convert those log occurrences into time-series metrics using Log-Based Metrics:
- Counter Metrics: Counts the number of log entries matching a filter over time.
- Distribution Metrics: Extracts numerical values from log payloads (e.g., latency values) to calculate latencies, P99s, or payload sizes.
# Create a Log-Based Counter Metric tracking database connection failures
gcloud logging metrics create db_conn_failures \
--description="Counts database connection pool failures" \
--log-filter="textPayload:\"Connection pool exhausted\""
Topic 4: Cloud Monitoring Workspace & Alert Policies
Cloud Monitoring aggregates metrics from GCE, GKE, Cloud SQL, and custom log-based metrics:
- Alerting Policies: Define Conditions (e.g.,
db_conn_failures > 5for 5 minutes), Duration, and Notification Channels (Slack, PagerDuty, Email, Pub/Sub webhook). - Cloud Trace: Distributed APM tracing that measures latency across microservice HTTP calls (automatically integrated into App Engine and GKE).
- Error Reporting: Aggregates stack traces from Go, Python, Java, Node.js, and displays grouped crash frequency dashboards.
Common mistake: Leaving the _Default bucket’s 30-day retention in place and calling it a logging strategy. Audit logs you will need for an investigation expire, high-volume application logs you will never read are retained at full price, and neither decision was made. Route with sinks — audit logs to a locked bucket in a separate project, noise to Cloud Storage or nowhere.