Everything in this lesson exists to answer one question: how do you reach something outside your subnet without sending the traffic across the public internet? The answers differ in cost, in scale, and in how badly they age.
Topic 1: The Four Mechanisms
Gateway endpoint — a route table entry, for S3 and DynamoDB only. No ENI, no security group, no hourly charge, no data processing charge. Traffic to the service’s prefix list is routed onto the AWS network directly. It cannot be reached from on-premises over Direct Connect or VPN, because on-prem traffic does not consult your subnet’s route table.
Interface endpoint (PrivateLink) — an ENI with a private IP inside your subnet, for almost every AWS service and for third-party or your own services. It has a security group, and it costs per hour per AZ plus per GB. Private DNS makes the service’s normal hostname resolve to the endpoint, so no client configuration changes.
VPC peering — a direct, non-transitive link between exactly two VPCs, any account, any region. No bandwidth bottleneck and no hourly charge; you pay only for cross-AZ or cross-region data transfer.
Transit Gateway — a regional hub. Each VPC, VPN and Direct Connect gateway attaches once; the TGW routes between attachments according to its own route tables. Hourly per attachment plus per GB processed.
Topic 2: Choosing Between Gateway and Interface Endpoints
| Gateway endpoint | Interface endpoint | |
|---|---|---|
| Services | S3, DynamoDB | ~150 AWS services, plus PrivateLink partners and your own |
| Mechanism | Route table + prefix list | ENI with a private IP |
| Cost | Free | ~$0.01/hour per AZ + per-GB processing |
| Security control | Endpoint policy | Endpoint policy and a security group |
| Reachable from on-prem | No | Yes, over DX or VPN |
| DNS | Unchanged; routing does the work | Private DNS overrides the public name |
Always add the S3 and DynamoDB gateway endpoints. They are free, they cut NAT data processing charges, and they remove a dependency on the NAT gateway for service traffic. There is no scenario where not having them is better.
Interface endpoints are a deliberate spend. The set that earns its keep in most accounts:
ssm, ssmmessages, ec2messages Session Manager without a NAT gateway
ecr.api, ecr.dkr (+ S3 gateway) container pulls off the NAT path
logs CloudWatch Logs from private subnets
secretsmanager, kms credentials without internet egress
sts role assumption in a fully private subnet
Two operational details that cause most interface-endpoint tickets:
- The security group. A new endpoint’s security group must allow inbound 443 from your instances’ CIDRs or security groups. Forget it and every call to that service hangs until it times out — no error, no rejection, just silence.
- One endpoint per AZ. Create the endpoint in each AZ’s subnet, or instances will cross AZs to reach it and you will pay inter-AZ transfer on every API call.
Endpoint policies are the underrated part. An endpoint policy restricts what can be done through that endpoint, regardless of IAM:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": "*",
"Action": ["s3:GetObject", "s3:PutObject", "s3:ListBucket"],
"Resource": ["arn:aws:s3:::my-app-data", "arn:aws:s3:::my-app-data/*"],
"Condition": { "StringEquals": { "aws:PrincipalOrgID": "o-abc123example" } }
}]
}
That is a data-exfiltration control: a compromised instance in this subnet cannot copy your data to an attacker’s bucket through this endpoint, because the endpoint will only talk to buckets you named. Pair it with a data subnet that has no 0.0.0.0/0 route and there is no path out at all.
Topic 3: Peering and Why It Stops Scaling
Peering is simple and, within its limits, excellent: no bandwidth ceiling, no single point of failure, no hourly cost.
Its three limits are absolute:
- Not transitive. A–B and B–C does not give you A–C. There is no configuration that changes this; it is the design.
- No overlapping CIDRs. The peering request is rejected outright.
- n(n−1)/2 connections, each needing route table entries on both sides. Four VPCs is 6 connections; eight is 28. Every new VPC means editing every existing route table.
The local route is also not shared across a peering connection: you must add explicit routes on both sides, and security groups must permit the other VPC’s traffic. Security group referencing across a peering connection works, but only within a region and only if you enable it.
Peering remains the right answer for exactly two VPCs that will stay two — a shared-services VPC and one consumer, say. The moment a third arrives, do the arithmetic before adding a fourth peering connection.
Topic 4: Transit Gateway and Route Table Segmentation
A Transit Gateway is a router. Attachments plug into it; TGW route tables decide which attachments can reach which others. That segmentation is the actual product — without it, a TGW is just a more expensive full mesh.
The standard segmented design:
TGW route table: PROD
attachments: prod-vpc-a, prod-vpc-b, shared-services, on-prem-vpn
routes: prod CIDRs, shared CIDRs, on-prem CIDRs
(no route to dev — dev is unreachable, not merely denied)
TGW route table: DEV
attachments: dev-vpc-a, dev-vpc-b, shared-services
routes: dev CIDRs, shared CIDRs
(no on-prem, no prod)
TGW route table: SHARED
attachments: shared-services
routes: everything — it must answer everyone
Each attachment is associated with one route table (which decides what it can reach) and can propagate its routes into several (which decides who can reach it). Getting that pair right is most of TGW operations, and misreading it is why “the route exists but traffic does not flow”.
Facts that shape designs:
- Regional. Cross-region needs TGW peering, and TGW peering is not transitive either.
- 50 Gbps per VPC attachment, aggregate across the attachment.
- An attachment lands in specific subnets — one per AZ. If an AZ has no attachment subnet, instances in that AZ reach the TGW across AZs and you pay for it.
- Appliance mode keeps a flow pinned to one AZ’s appliance for stateful inspection. Without it, asymmetric routing breaks any stateful firewall you insert.
Cost is the honest trade-off: per-attachment hourly plus per-GB processing, so a chatty pair of VPCs is cheaper peered. The usual mature design is a TGW for the general mesh, plus direct peering for the one or two very high-volume paths.
Topic 5: Hybrid — VPN and Direct Connect
| Site-to-Site VPN | Direct Connect | |
|---|---|---|
| Medium | IPsec over the internet | Dedicated physical circuit |
| Setup | Minutes | Weeks to months |
| Bandwidth | ~1.25 Gbps per tunnel | 1/10/100 Gbps |
| Latency | Variable — it is the internet | Consistent |
| Cost | Cheap hourly + data | Port fee + much cheaper egress |
| Encryption | Always | None by default — run a VPN over it or use MACsec |
Each VPN connection gives you two tunnels to two AWS endpoints for redundancy, and a surprising number of production VPNs run with only one configured. Configure both.
That last row of the table is the one that catches people: Direct Connect is a private circuit, not an encrypted one. Compliance regimes that require encryption in transit are not satisfied by “it is a dedicated line”.
The resilient hybrid pattern is DX as primary with a VPN as backup over the internet, both attached to the same Transit Gateway, with BGP preferring DX. When the circuit fails, BGP converges to the VPN and you take a bandwidth hit rather than an outage. Test the failover on purpose, on a schedule — an untested backup path is a hypothesis.
Topic 6: PrivateLink for Your Own Services
The same machinery that exposes AWS services privately can expose yours. Put a Network Load Balancer in front of your service, create a VPC endpoint service from it, and consumers create an interface endpoint in their own VPC.
Why this beats peering for a service:
- No CIDR coordination. The consumer’s address space is irrelevant, and overlap does not matter.
- One direction only. The consumer can reach your service; you cannot reach into their VPC. Peering is bidirectional by nature.
- No route table changes on either side, and no blast radius growth as consumers are added.
- You approve each consumer explicitly and can revoke one without affecting the others.
This is how SaaS vendors expose products inside your VPC, and it is the right internal pattern too — a platform team offering a service to twenty product teams should offer an endpoint service, not twenty peering connections.
Try it yourself: add the S3 gateway endpoint, then compare NAT gateway BytesOutToDestination in CloudWatch before and after. On a log-shipping workload the drop is dramatic and immediate, and it is the easiest cost win in this entire module.
Common mistake: creating an interface endpoint with private DNS enabled while a Route 53 private hosted zone for the same service name already exists. Both try to answer, one wins non-deterministically per resolver, and you get intermittent failures that correlate with nothing. Pick one mechanism per service name and delete the other.