This final capstone lesson connects GCP’s managed messaging and data analytics engine with enterprise security perimeters.
Topic 1: Cloud Pub/Sub vs. Cloud Tasks
GCP provides two primary asynchronous messaging services that sound similar but serve distinct architecture patterns:
| Feature | Cloud Pub/Sub | Cloud Tasks |
|---|---|---|
| Messaging Model | Publisher-Subscriber (1-to-many fan-out) | Targeted Queue (1-to-1 execution) |
| Publisher Control | Publisher is decoupled; knows nothing about subscribers | Publisher stays in full control of target execution endpoint |
| Features | High-throughput streaming, global message ingestion | Rate limiting, deduplication, scheduled delivery times, retries |
| Delivery Guarantee | At-least-once delivery (unordered by default) | Configurable rate/retry limits per queue |
| Typical Use Case | Real-time event streaming, log sinks, telemetry | Asynchronous background tasks, transactional emails, webhooks |
# Create a Pub/Sub Topic and a Push/Pull Subscription
gcloud pubsub topics create telemetry-events
gcloud pubsub subscriptions create sub-analytics \
--topic=telemetry-events \
--ack-deadline=30
Topic 2: BigQuery Cost & Query Optimization (Partitioning vs. Clustering)
BigQuery is Google’s serverless, petabyte-scale columnar data warehouse. On-demand BigQuery queries are billed based on the number of bytes scanned ($5 per TB). Running SELECT * on an unpartitioned 50 TB table costs $250 for a single query!
Optimization Strategy 1: Partitioned Tables:
Divides large tables into segments based on a date column (_PARTITIONDATE or timestamp). When a query includes WHERE timestamp >= '2026-08-01', BigQuery scans ONLY the matching date partition files, cutting query costs by 95%+.
Optimization Strategy 2: Clustered Tables:
Sorts data automatically based on up to 4 high-cardinality columns (e.g., customer_id, order_status). Queries filtering by customer_id skip non-matching data blocks within partitions.
-- Create a Partitioned and Clustered Table in BigQuery SQL
CREATE TABLE `my_project.analytics.orders`
(
order_id STRING,
customer_id STRING,
amount NUMERIC,
order_timestamp TIMESTAMP
)
PARTITION BY DATE(order_timestamp)
CLUSTER BY customer_id, order_id;
Topic 3: Managed Database Decision Matrix
GCP offers specialized database engines tailored for specific workload characteristics:
- Cloud SQL: Managed MySQL, PostgreSQL, and SQL Server. Regional availability, max 30 TB per instance. Ideal for standard monolithic or microservice relational workloads.
- Cloud Spanner: Enterprise globally distributed relational database. Provides Strong Consistency (External Consistency / TrueTime) with horizontal scaling across regions and continents.
- Cloud Bigtable: Managed NoSQL wide-column database (HBase API). Ultra-low single-digit millisecond latency for massive write-heavy workloads (IoT sensors, ad tech, financial ticker data).
- Firestore / Datastore: Serverless document NoSQL database for web and mobile state, product catalogs, and user profiles.
- Memorystore: Managed in-memory Redis / Memcached for high-speed caching.
Topic 4: Zero-Trust Security: IAP, Cloud Armor & VPC Service Controls
1. Identity-Aware Proxy (IAP)
- Implements a Zero-Trust Security Model. Controls access to HTTPS applications and SSH/RDP to GCE VMs based on user identity and context (device posture, IP location) without requiring a VPN or exposing public IPs.
- TCP Forwarding via IAP allows secure SSH access to private VMs (
gcloud compute ssh --tunnel-through-iap).
2. Cloud Armor
- Enterprise DDoS protection and Web Application Firewall (WAF) enforced at Google’s global Edge Points of Presence (PoPs).
- Filters malicious traffic, rate-limits IP blocks, and blocks OWASP Top 10 attacks before traffic reaches your VPC.
3. VPC Service Controls (VPC-SC)
- Establishes a security perimeter around sensitive GCP service APIs (GCS, BigQuery, Cloud SQL).
- Prevents data exfiltration: Even if an attacker gains valid credentials, VPC Service Controls blocks copying or reading data from outside authorized networks or perimeter boundaries.
Common mistake: Running SELECT * against a partitioned table without a partition filter. BigQuery bills on bytes scanned, so the query that cost nothing in development scans the entire table in production — and the fix, require_partition_filter on the table, turns the mistake into an error rather than an invoice.