Pub/Sub In Depth: Delivery, Ordering and Backlogs

Push versus pull, ack deadlines and the redelivery loop, ordering keys and what they cost, dead-letter topics, exactly-once, and the one metric that tells you a consumer is broken.

advanced 22 min lesson hands-on task included

Pub/Sub is easy to start with and has a small number of behaviours that decide whether a system built on it works under load. This lesson is those behaviours.


Topic 1: Delivery, and the Loop That Governs Everything

ONE TOPIC, MANY SUBSCRIPTIONS — EACH GETS ITS OWN COPY topic publishers write here pull subscription worker asks for messages push subscription Pub/Sub POSTs to your URL THE ACK DEADLINE IS THE TRAP Not acked in time → redelivered. A slow consumer therefore gets the same message again while still processing the first copy. Extend the deadline, or make the handler idempotent. Ideally both. DEAD-LETTER TOPIC After N delivery attempts, the message moves aside instead of retrying forever and blocking the backlog. --max-delivery-attempts=5 --dead-letter-topic=… ORDERING COSTS THROUGHPUT An ordering key serialises delivery for that key — correct when order matters, and a bottleneck when it does not. Exactly-once is available on pull, in one region, with a cost. THE METRIC THAT MATTERS IS AGE, NOT DEPTH subscription/oldest_unacked_message_age — a burst that drains in a second is fine; one message stuck for an hour is not. Retention holds unacked messages up to 7 days, and a subscription with no consumer expires after 31 days of inactivity by default.
The ack deadline is the centre of the system. A message not acked in time is redelivered — which is why every consumer must be idempotent, and why slow processing looks like duplication.
publish → message stored (redundantly, up to 7 days by default)
        → delivered to each subscription independently
        → consumer has ack_deadline seconds to ack
             ack      → message removed from that subscription
             nack     → immediate redelivery
             silence  → redelivered when the deadline expires

Every subscription gets its own copy. Two subscriptions on one topic are two independent backlogs with independent acks — which is how you fan out to a real-time consumer and a BigQuery archive without them affecting each other.

The ack deadline is the parameter people get wrong. Set it to 10 seconds and a consumer that takes 30 seconds to process will see every message three times, do the work three times, and never make progress:

gcloud pubsub subscriptions create orders-worker --topic=orders \
  --ack-deadline=60 \
  --message-retention-duration=7d \
  --dead-letter-topic=orders-dlq --max-delivery-attempts=5

The client libraries extend the deadline automatically while a message is being processed, up to maxAckExtensionPeriod. That is the right mechanism for variable work — set a modest deadline and let the library hold it — but the extension cap is a real ceiling, and work that can exceed it belongs in a job triggered by the message rather than in the handler.

Delivery is at-least-once. Idempotency is not optional, and the pattern is the same one as bucket notifications: dedupe on a stable key, in a store, before doing the work.


Topic 2: Push, Pull, and BigQuery Subscriptions

TypeHow it worksChoose it when
PullConsumer opens a streaming connection and requests messagesYou want flow control, high throughput, or long processing
PushPub/Sub POSTs to an HTTPS endpointThe consumer is Cloud Run/Functions and you want scale-to-zero
BigQueryPub/Sub writes rows directlyYou just want the data in a table with no code
Cloud StoragePub/Sub writes batched filesCheap archival, no consumer to run

Push, done properly, is authenticated:

gcloud pubsub subscriptions create orders-push --topic=orders \
  --push-endpoint=https://processor-xyz.a.run.app/events \
  --push-auth-service-account=pubsub-push@acme.iam.gserviceaccount.com

Pub/Sub mints an OIDC token for that service account; Cloud Run verifies it. The endpoint must not be public — grant roles/run.invoker to the push service account only. A push endpoint with allUsers invoker is an open ingestion point for anyone who finds the URL.

Push backs off on failure, which is a feature and a trap: a consumer returning 500 slows delivery, so a partial outage becomes a growing backlog rather than a flood. That is correct behaviour, and it means the backlog metric is your outage signal.

BigQuery subscriptions remove a whole service from the architecture:

gcloud pubsub subscriptions create orders-to-bq --topic=orders \
  --bigquery-table=acme:events.orders \
  --use-topic-schema --drop-unknown-fields

With a topic schema, Pub/Sub validates at publish time and writes typed rows. No Dataflow job, no consumer to operate. Use it whenever the requirement really is “put these events in a table” — which it often is.


Topic 3: Ordering, and What It Costs

By default there is no ordering. With an ordering key, messages sharing that key are delivered in publish order:

publisher = pubsub_v1.PublisherClient(
    publisher_options=pubsub_v1.types.PublisherOptions(enable_message_ordering=True)
)
publisher.publish(topic, data, ordering_key=f"account-{account_id}")
gcloud pubsub subscriptions create orders-ordered --topic=orders \
  --enable-message-ordering

The trade, stated plainly:

  • Throughput per key is serial. Parallelism comes from having many keys, so account-1234 is a good key and orders is a terrible one.
  • A stuck message blocks its key. Everything behind it waits until it is acked or dead-lettered — by design, and exactly why the dead-letter topic matters more here.
  • Publishing is slower, because ordered publishes cannot be batched as freely.
  • A publish error on a key blocks that key until you call resume_publish.

Order by account, tenant, device or entity — never by something coarse. And ask first whether you need it: many systems that reach for ordering actually need idempotency plus a version number on the record, which has none of these costs.


Topic 4: Dead-Letter Topics and Retries

gcloud pubsub subscriptions update orders-worker \
  --dead-letter-topic=orders-dlq \
  --max-delivery-attempts=5 \
  --min-retry-delay=10s --max-retry-delay=600s

Two grants are required, and missing either means dead-lettering silently does not happen:

SA="serviceAccount:service-${PROJECT_NUMBER}@gcp-sa-pubsub.iam.gserviceaccount.com"
gcloud pubsub topics add-iam-policy-binding orders-dlq \
  --member="$SA" --role=roles/pubsub.publisher
gcloud pubsub subscriptions add-iam-policy-binding orders-worker \
  --member="$SA" --role=roles/pubsub.subscriber

Give the dead-letter topic its own subscription, or the messages you carefully rescued expire after 7 days having been read by nobody.

Exponential backoff (--min-retry-delay / --max-retry-delay) is what stops a failing consumer being hammered. Without it, a message that fails instantly is retried instantly, and the retry loop becomes a load test against your own broken service.

Alert on dead_letter_message_count > 0. A message that failed five times is essentially never transient, and it is the highest-signal, lowest-noise alert in a Pub/Sub system.


Topic 5: Exactly-Once, Flow Control and Replay

Exactly-once delivery is available on regional subscriptions, and is narrower than the name suggests:

gcloud pubsub subscriptions create orders-eos --topic=orders \
  --enable-exactly-once-delivery --message-storage-policy-allowed-regions=europe-west1

It guarantees the message is not redelivered after a successful ack. It does not make your side effects atomic: if you write to a database and then crash before acking, the message returns. You still need idempotency; exactly-once removes a class of duplicates, not the need for the pattern.

Flow control protects the consumer, and is client-side:

flow_control = pubsub_v1.types.FlowControl(max_messages=100, max_bytes=10*1024*1024)
subscriber.subscribe(subscription, callback=handle, flow_control=flow_control)

Without it, the client pulls as fast as it can and a memory-bound consumer falls over under a burst.

Replay is the underused feature. Seek moves a subscription’s position:

# Reprocess the last two hours
gcloud pubsub subscriptions seek orders-worker \
  --time=$(date -u -d '2 hours ago' +%Y-%m-%dT%H:%M:%SZ)

# Discard a backlog entirely
gcloud pubsub subscriptions seek orders-worker --time=$(date -u +%Y-%m-%dT%H:%M:%SZ)

Seeking backwards requires the messages to still be within message-retention-duration, and --retain-acked-messages is what makes replay possible after acking — worth enabling on subscriptions whose consumers you expect to have bugs, which is all of them.

Snapshots capture a subscription’s ack state so you can deploy a risky consumer, and seek back to the snapshot if it processed everything wrongly. That is a genuinely good deployment practice for a stateful consumer.


Topic 6: Operating It

The metrics that matter, in order:

subscription/oldest_unacked_message_age    ← the health signal
subscription/num_undelivered_messages      ← depth; interpret with the above
subscription/dead_letter_message_count     ← alert on any
subscription/ack_message_count             ← throughput
topic/send_request_count by response_code  ← publish-side errors
subscription/push_request_count by code    ← push consumer health

Alert on age, not on depth. A backlog of 100,000 messages that drains in 30 seconds is a healthy burst. One message stuck for two hours is a broken consumer, and a depth alarm misses it entirely.

A workable starting point: warn at 5 minutes of message age, page at 30, and tune from what your consumer’s normal p99 actually is.

Quotas and limits you will meet:

message size            10 MB          → put the payload in GCS, publish the pointer
retention               7 days default, 31 days maximum
attributes              100 per message, 1024 bytes per key/value
publish throughput      high, but per-region quota — check before a launch
ordering key throughput 1 MB/s per key  ← the practical ordering limit

Schemas are worth the small effort:

gcloud pubsub schemas create order-v1 --type=avro --definition-file=order.avsc
gcloud pubsub topics create orders --schema=order-v1 --message-encoding=json

A schema turns “the producer changed a field and three consumers broke” into a publish-time rejection, and it is the prerequisite for a BigQuery subscription with real column types.

Try it yourself: stop your consumer for ten minutes with traffic flowing, then start it again and watch oldest_unacked_message_age spike and recover. That shape is what a real incident looks like on this metric, and seeing it once makes the threshold choice obvious.

Common mistake: setting the ack deadline to the default 10 seconds for a handler that takes 30. Every message is redelivered while still being processed, the work is done repeatedly, throughput collapses under retries, and the system looks like it is under load when it is actually fighting itself. Measure the handler’s p99 and set the deadline above it — or let the client library extend it and cap the extension deliberately.