Bucket Notifications and Eventarc

Turning an object change into a message, why every consumer must be idempotent, the loop that costs money, and choosing between Pub/Sub notifications and Eventarc.

intermediate 20 min lesson hands-on task included

Object storage becomes an event source the moment you attach a notification, and that is how most data pipelines on GCP actually start. The mechanics are simple; the delivery semantics are where the bugs live.


Topic 1: The Pipeline

AN OBJECT CHANGE BECOMES A MESSAGE — THEN ANYTHING CAN CONSUME IT bucket event FINALIZE · DELETE ARCHIVE · METADATA_UPDATE Pub/Sub topic attributes carry the object name consumer Cloud Run · Functions · Dataflow TWO WAYS TO WIRE IT gcloud storage buckets notifications create gs://b \ --topic=uploads --event-types=OBJECT_FINALIZE …or Eventarc, which adds filtering and CloudEvents format. DELIVERY IS AT-LEAST-ONCE The same object event can arrive twice, and out of order. Make the consumer idempotent — key on generation number, not on the object name. THE LOOP THAT COSTS MONEY A consumer that writes back into the same bucket triggers itself. Write to a different bucket, or filter by prefix — and set a dead-letter topic. The GCS service agent needs roles/pubsub.publisher on the topic — the notification silently produces nothing without it.
An object change becomes a Pub/Sub message and anything can consume it. The two lower panels are the parts that cause incidents — at-least-once delivery, and the consumer that triggers itself.
gcloud storage buckets notifications create gs://acme-uploads \
  --topic=uploads \
  --event-types=OBJECT_FINALIZE \
  --payload-format=json \
  --object-prefix=incoming/

The event types:

TypeFires when
OBJECT_FINALIZEA new object is created, or an existing one overwritten
OBJECT_DELETEAn object is deleted, or overwritten (the old version)
OBJECT_ARCHIVEA versioned object becomes noncurrent
OBJECT_METADATA_UPDATEMetadata changes

Subscribe to the narrowest set you need. A pipeline that only cares about uploads should not receive delete events, and --object-prefix keeps it to the paths that matter — both reduce cost and reduce the chance of a surprising trigger.

The message carries attributes, not the object. bucketId, objectId, eventType, objectGeneration are attributes; the payload is the object’s metadata. Your consumer reads the object itself if it needs the bytes.


Topic 2: Delivery Semantics, and the Bug They Cause

Delivery is at-least-once and unordered. The same event can arrive twice, and events for two objects can arrive in either order.

That produces a specific class of bug: a consumer that appends a row per event double-counts, or a consumer that assumes “create before update” processes them backwards.

Make the handler idempotent, keyed on generation:

def handle(event):
    bucket = event["attributes"]["bucketId"]
    name   = event["attributes"]["objectId"]
    gen    = event["attributes"]["objectGeneration"]   # unique per object version

    if already_processed(bucket, name, gen):
        return  # ack and move on
    process(bucket, name, gen)
    mark_processed(bucket, name, gen)

Key on objectGeneration, not on the object name. An object overwritten twice produces two events with the same name and different generations — deduplicating on the name loses the second write.

And handle the object being gone. By the time your consumer reads it, a later delete may already have happened. A 404 is a normal outcome, not an error to retry forever.


Topic 3: The Loop That Costs Money

A consumer that writes its output back into the same bucket triggers itself. Each output triggers another invocation, which produces another output.

upload → event → consumer writes thumbnail to same bucket
              → event → consumer writes thumbnail of thumbnail → …

Three ways to prevent it, in order of robustness:

  • Write to a different bucket. Simplest, and it cannot regress.
  • Use --object-prefix on the notification so only incoming/ triggers it and the consumer writes to processed/.
  • Check the object path in the handler and return early. Works, and depends on nobody changing the paths later.

Always set a dead-letter topic, so a message that fails repeatedly stops being retried forever:

gcloud pubsub subscriptions create uploads-worker \
  --topic=uploads --dead-letter-topic=uploads-dlq --max-delivery-attempts=5 \
  --ack-deadline=60

Without one, a poison message retries indefinitely, consumes quota, and blocks nothing visibly — the backlog just never drains.


Topic 4: The Permission Everyone Misses

The GCS service agent publishes the messages, and it needs permission on the topic:

SERVICE_ACCOUNT=$(gcloud storage service-agent --project=acme)
gcloud pubsub topics add-iam-policy-binding uploads \
  --member="serviceAccount:${SERVICE_ACCOUNT}" --role=roles/pubsub.publisher

Without it, the notification is created successfully and silently produces nothing. The bucket reports the notification exists; no message ever arrives; there is no error anywhere obvious. It is the single most common failure with this feature.

# Verify a notification is actually configured
gcloud storage buckets notifications list gs://acme-uploads

Topic 5: Eventarc, and When to Prefer It

Eventarc is the general eventing layer: it delivers events from around a hundred Google sources — including Cloud Storage — to Cloud Run, Workflows or GKE, in CloudEvents format.

gcloud eventarc triggers create process-upload \
  --destination-run-service=processor \
  --destination-run-region=europe-west1 \
  --event-filters="type=google.cloud.storage.object.v1.finalized" \
  --event-filters="bucket=acme-uploads" \
  --service-account=eventarc@acme.iam.gserviceaccount.com
GCS → Pub/Sub notificationEventarc
SourcesCloud Storage only~100 Google services, plus Audit Log events
FormatGoogle’s own JSONCloudEvents, portable
FilteringEvent type and prefixRicher attribute filters
Delivery targetAnything that reads Pub/SubCloud Run, Workflows, GKE
UnderneathA Pub/Sub topic you ownA Pub/Sub topic Eventarc manages

Prefer Eventarc when you want one eventing model across services, CloudEvents portability, or events from an Audit Log (for example “someone changed a firewall rule”). Prefer the direct notification when you want to own the topic and fan out to several unrelated consumers with different retention.

Both are Pub/Sub underneath, so everything in the next lesson about ack deadlines, dead-letter topics and backlog age applies to either.


Topic 6: Operating the Pipeline

The metrics that tell you it is healthy — and the first is not the obvious one:

subscription/oldest_unacked_message_age   ← the one that matters
subscription/num_undelivered_messages
subscription/dead_letter_message_count    ← alert on any non-zero value
run.googleapis.com/request_count by response code

Alert on message age, not backlog depth. Ten thousand messages that drain in a second is a healthy burst; one message stuck for an hour is a broken consumer. Depth alarms fire on the first and miss the second.

A non-zero dead-letter count is always worth a page in a pipeline that is supposed to be reliable — it means a message failed five times, which is rarely transient.

Test the failure paths deliberately:

□ upload the same object twice — is the outcome identical?
□ upload a malformed object — does it dead-letter rather than retry forever?
□ delete the object mid-processing — does the 404 ack rather than loop?
□ stop the consumer for an hour — does the backlog drain afterwards?
□ check retention: unacked messages are kept 7 days by default, then dropped

Try it yourself: remove the GCS service agent’s publisher role on the topic and upload an object. Nothing arrives, nothing errors, and the notification still shows as configured. Reproducing that silence once makes it instantly recognisable later.

Common mistake: building the consumer to assume exactly-once, in-order delivery because that is what the happy path looks like in testing. In production the same upload event arrives twice during a redeploy, the handler processes it twice, and the resulting duplicate rows or double-charged transactions are traced back to the pipeline days later. Idempotency keyed on generation costs a few lines and removes the entire class.