The Log Router: Sinks, Filters and Exports

How routing actually evaluates, the four destinations and what each is for, aggregated sinks at the organisation level, and the writer identity that makes exports silently fail.

advanced 20 min lesson hands-on task included

The Log Router sits between every log entry and everywhere it can go. It is a small piece of machinery, and understanding its evaluation order is what separates a logging design from a pile of sinks.


Topic 1: How Routing Evaluates

ONE ROUTER, FOUR DESTINATIONS — CHOOSE BY WHAT YOU WILL DO WITH IT log bucket query in Logs Explorer operational debugging BigQuery SQL over log data analytics, DORA metrics Cloud Storage cheap, immutable, cold long retention, audit Pub/Sub stream to anything SIEM, real-time alerting AN AGGREGATED SINK IS THE ONE THAT SCALES gcloud logging sinks create org-audit \ bigquery.googleapis.com/... --organization=… \ --include-children THE PERMISSION EVERYONE FORGETS Each sink gets its own writer identity — a service account Google creates. It needs write access on the destination, or the sink silently delivers nothing. SINKS COPY — THEY DO NOT MOVE A sink to BigQuery still leaves the entry in _Default, and you pay for both. Pair the sink with an exclusion when the copy is the point. Route audit logs to a separate project nobody in the workload project can write to. That is the tamper-evidence.
Every entry is offered to every sink independently. That is the property to internalise — sinks do not consume entries, so one entry can reach four destinations and be billed for each.
entry arrives
  → for EACH sink in scope:
       does the sink's inclusion filter match?      no  → skip this sink
       does any exclusion filter match?             yes → skip this sink
       → write to that sink's destination

Three consequences that explain most surprises:

  • Sinks are independent, not a chain. There is no “first match wins”. An entry matching three sinks goes to three destinations.
  • _Default is a sink too. Adding a custom sink does not remove an entry from _Default; you exclude it there if you do not want the duplicate.
  • Exclusions live on a sink. An exclusion on your BigQuery sink does not stop the entry reaching _Default. To stop ingestion entirely, exclude on _Default as well.
gcloud logging sinks list
gcloud logging sinks describe security-to-bq --format='yaml(filter, destination, exclusions)'

Topic 2: The Four Destinations

DestinationUse it forQuery surfaceCost shape
Log bucketAnything you will read in Logs ExplorerFilter language, plus SQL with AnalyticsIngestion, then storage past retention
BigQueryAnalysis, joins to business data, long retentionFull SQLStorage + query bytes scanned
Cloud StorageCheap, immutable, compliance archiveNone until you load itStorage class; very cheap on Archive
Pub/SubReal-time reaction — alerting, SIEM, custom processingNone; it is a streamPer message + subscription backlog

Log bucket is the default answer. Reach for the others when you have a reason.

BigQuery — use partitioned tables or the cost grows without bound:

gcloud logging sinks create app-to-bq \
  bigquery.googleapis.com/projects/acme/datasets/app_logs \
  --log-filter='resource.type="cloud_run_revision" severity>=WARNING' \
  --use-partitioned-tables

Without --use-partitioned-tables you get date-suffixed tables and every query scans more than it needs. Set a partition expiry on the dataset so old data ages out on its own.

Cloud Storage — the compliance archive, and the one place bucket lock belongs:

gcloud logging sinks create audit-archive \
  storage.googleapis.com/acme-audit-archive \
  --log-filter='logName:"cloudaudit.googleapis.com"'

Objects arrive hourly as newline-delimited JSON under a date path. Pair with a lifecycle rule to Archive class and, where the requirement is real, a retention policy with bucket lock — which even an org admin cannot shorten. Note the latency: entries appear within the hour, not in seconds, so this is an archive and never an alerting path.

Pub/Sub — the real-time path, and the one to keep narrow:

gcloud logging sinks create critical-to-pubsub \
  pubsub.googleapis.com/projects/acme/topics/critical-logs \
  --log-filter='severity>=ERROR AND resource.labels.namespace_name="checkout"'

Keep the filter tight. A Pub/Sub sink on everything is an expensive firehose that no consumer keeps up with, and the backlog age metric will tell you so within a day.


Topic 3: The Writer Identity — the Silent Failure

Every sink writes as a service account Google creates for it, and that account starts with no permission on your destination.

# 1. Create the sink and read back its writer identity
gcloud logging sinks create app-to-bq \
  bigquery.googleapis.com/projects/acme/datasets/app_logs \
  --log-filter='severity>=WARNING'

WRITER=$(gcloud logging sinks describe app-to-bq --format='value(writerIdentity)')

# 2. Grant it on the destination — this step is the one that gets missed
gcloud projects add-iam-policy-binding acme \
  --member="$WRITER" --role=roles/bigquery.dataEditor

The grants per destination:

log bucket      roles/logging.bucketWriter
BigQuery        roles/bigquery.dataEditor    (on the dataset or project)
Cloud Storage   roles/storage.objectCreator  (on the bucket)
Pub/Sub         roles/pubsub.publisher       (on the topic)

Without the grant, the sink exists, reports no error, and delivers nothing. This is the same failure shape as the GCS notification permission in the previous stage, and it is worth recognising as a pattern: on GCP, a service-to-service write that lacks IAM usually fails silently at the producer.

Verify by observation, not by configuration. The only reliable check is that data arrived:

bq query --nouse_legacy_sql \
  'SELECT COUNT(*) FROM `acme.app_logs.INFORMATION_SCHEMA.TABLES`'

gcloud storage ls gs://acme-audit-archive/ --recursive | head

Sinks recreated with --uniq-writer-identity get a fresh identity, which means the grant has to be redone — a sink that worked and then stopped after being recreated is almost always this.


Topic 4: Aggregated Sinks

A sink at the folder or organisation level with --include-children captures logs from every project underneath, including projects created later:

gcloud logging sinks create org-audit-archive \
  storage.googleapis.com/acme-org-audit \
  --organization=123456789012 \
  --include-children \
  --log-filter='logName:"cloudaudit.googleapis.com%2Factivity"'

This is the only way to build a logging design that does not depend on every team configuring it. A new project inherits the org sink the day it is created, with nothing for its owners to do and nothing for them to accidentally remove.

The pairing that makes it a control rather than a suggestion: an org policy or a deny rule preventing project owners from deleting the org sink’s destination, plus bucket lock on the archive. Otherwise the audit trail is only as durable as the least careful project owner.

Filter with resource.labels.project_id to route different projects to different destinations from one aggregated sink — for example, sending regulated projects’ logs to a bucket in a specific region.


Topic 5: A Design That Holds Up

The four-destination shape most mature setups converge on:

ORG LEVEL (aggregated, include-children)
  audit activity  → GCS archive, Archive class, bucket lock, 7 years
  audit activity  → log bucket in the security project, 400 days, Log Analytics

PROJECT LEVEL
  app logs        → _Default, 30 days, health checks excluded
  app WARNING+    → BigQuery, partitioned, 90-day partition expiry
  ERROR in prod   → Pub/Sub → alerting consumer

Read against the questions from the previous lesson: what is kept, for how long, who can read it, what it costs. Each line answers all four, which is the point of writing the design down rather than accumulating sinks.

Manage it in Terraform, because the writer-identity grant belongs next to the sink:

resource "google_logging_project_sink" "app_to_bq" {
  name                   = "app-to-bq"
  destination            = "bigquery.googleapis.com/projects/acme/datasets/app_logs"
  filter                 = "severity>=WARNING"
  unique_writer_identity = true

  bigquery_options { use_partitioned_tables = true }
}

resource "google_project_iam_member" "app_to_bq_writer" {
  project = "acme"
  role    = "roles/bigquery.dataEditor"
  member  = google_logging_project_sink.app_to_bq.writer_identity
}

The dependency between those two resources is exactly the step people forget by hand, which is a good reason for this to be code.


Topic 6: Log-Based Metrics

The router’s other output: turning matching entries into a metric you can alert on.

gcloud logging metrics create payment_failures \
  --description="Failed payment attempts" \
  --log-filter='resource.type="cloud_run_revision"
                jsonPayload.event="payment_failed"'

Distribution metrics extract a numeric value, which is how you get a latency histogram out of structured logs without instrumenting the application:

gcloud logging metrics create checkout_latency \
  --log-filter='jsonPayload.duration_ms:*' \
  --value-extractor='EXTRACT(jsonPayload.duration_ms)' \
  --metric-kind=DELTA --value-type=DISTRIBUTION

Two limits to know before you build an alerting strategy on these:

  • They only count entries ingested after the metric was created. There is no backfill, so create the metric before you need the history.
  • They are subject to the same exclusions. An entry excluded from _Default never produces the metric — which is a genuine trap when someone excludes a noisy log type that a metric was counting.

Prefer a real application metric where you can. A log-based metric is the right tool when you cannot change the application, and a slightly awkward one when you can.

Try it yourself: create a sink to BigQuery, skip the writer-identity grant, and wait ten minutes. No error, no data, and a sink that describes as healthy. Doing this deliberately once is what makes you check writerIdentity first the next time an export “stopped working”.

Common mistake: routing everything to BigQuery “so we have it”, with no partitioning and no partition expiry. Storage grows without bound, every ad-hoc query scans months of data, and the logging bill arrives as a BigQuery bill so nobody connects the two. Filter at the sink, partition the tables, and expire the partitions.