Endpoints, Peering and Transit Gateway

Reaching AWS services and other VPCs without touching the internet — gateway versus interface endpoints, why peering does not scale past three VPCs, and what a Transit Gateway route table is really for.

intermediate 22 min lesson hands-on task included

Everything in this lesson exists to answer one question: how do you reach something outside your subnet without sending the traffic across the public internet? The answers differ in cost, in scale, and in how badly they age.


Topic 1: The Four Mechanisms

FOUR WAYS TO REACH SOMETHING WITHOUT TOUCHING THE INTERNET GATEWAY ENDPOINT — a route table entry subnet prefix list S3 Only S3 and DynamoDB. No ENI, no hourly cost, no data-processing charge — and it kills NAT bytes. Not reachable from on-prem over DX/VPN. INTERFACE ENDPOINT — an ENI in your subnet subnet ENI + private DNS SSM Almost every service. Costs per hour per AZ plus per GB — so an endpoint per AZ, or you pay cross-AZ. Has a security group. Forgetting it looks like a hang. VPC PEERING — NOT TRANSITIVE A B C A cannot reach C through B Overlapping CIDRs cannot peer at all. n² route tables to maintain. TRANSIT GATEWAY — A HUB THAT ROUTES TGW route tables prod shared dev on-prem Separate TGW route tables are how you keep dev out of prod. THE ORDER TO REACH FOR THEM Gateway endpoint if it is S3 or DynamoDB. Interface endpoint for an AWS API. Peering for exactly two VPCs that will stay two. Transit Gateway the moment there is a third — the crossover is around four VPCs, and it arrives sooner than anyone plans for.
Two endpoint types for AWS services, two topologies for VPC-to-VPC. The crossover from peering to Transit Gateway arrives at around four VPCs — and it always arrives sooner than the original design assumed.

Gateway endpoint — a route table entry, for S3 and DynamoDB only. No ENI, no security group, no hourly charge, no data processing charge. Traffic to the service’s prefix list is routed onto the AWS network directly. It cannot be reached from on-premises over Direct Connect or VPN, because on-prem traffic does not consult your subnet’s route table.

Interface endpoint (PrivateLink) — an ENI with a private IP inside your subnet, for almost every AWS service and for third-party or your own services. It has a security group, and it costs per hour per AZ plus per GB. Private DNS makes the service’s normal hostname resolve to the endpoint, so no client configuration changes.

VPC peering — a direct, non-transitive link between exactly two VPCs, any account, any region. No bandwidth bottleneck and no hourly charge; you pay only for cross-AZ or cross-region data transfer.

Transit Gateway — a regional hub. Each VPC, VPN and Direct Connect gateway attaches once; the TGW routes between attachments according to its own route tables. Hourly per attachment plus per GB processed.


Topic 2: Choosing Between Gateway and Interface Endpoints

Gateway endpointInterface endpoint
ServicesS3, DynamoDB~150 AWS services, plus PrivateLink partners and your own
MechanismRoute table + prefix listENI with a private IP
CostFree~$0.01/hour per AZ + per-GB processing
Security controlEndpoint policyEndpoint policy and a security group
Reachable from on-premNoYes, over DX or VPN
DNSUnchanged; routing does the workPrivate DNS overrides the public name

Always add the S3 and DynamoDB gateway endpoints. They are free, they cut NAT data processing charges, and they remove a dependency on the NAT gateway for service traffic. There is no scenario where not having them is better.

Interface endpoints are a deliberate spend. The set that earns its keep in most accounts:

ssm, ssmmessages, ec2messages     Session Manager without a NAT gateway
ecr.api, ecr.dkr (+ S3 gateway)   container pulls off the NAT path
logs                              CloudWatch Logs from private subnets
secretsmanager, kms               credentials without internet egress
sts                               role assumption in a fully private subnet

Two operational details that cause most interface-endpoint tickets:

  • The security group. A new endpoint’s security group must allow inbound 443 from your instances’ CIDRs or security groups. Forget it and every call to that service hangs until it times out — no error, no rejection, just silence.
  • One endpoint per AZ. Create the endpoint in each AZ’s subnet, or instances will cross AZs to reach it and you will pay inter-AZ transfer on every API call.

Endpoint policies are the underrated part. An endpoint policy restricts what can be done through that endpoint, regardless of IAM:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": "*",
    "Action": ["s3:GetObject", "s3:PutObject", "s3:ListBucket"],
    "Resource": ["arn:aws:s3:::my-app-data", "arn:aws:s3:::my-app-data/*"],
    "Condition": { "StringEquals": { "aws:PrincipalOrgID": "o-abc123example" } }
  }]
}

That is a data-exfiltration control: a compromised instance in this subnet cannot copy your data to an attacker’s bucket through this endpoint, because the endpoint will only talk to buckets you named. Pair it with a data subnet that has no 0.0.0.0/0 route and there is no path out at all.


Topic 3: Peering and Why It Stops Scaling

Peering is simple and, within its limits, excellent: no bandwidth ceiling, no single point of failure, no hourly cost.

Its three limits are absolute:

  1. Not transitive. A–B and B–C does not give you A–C. There is no configuration that changes this; it is the design.
  2. No overlapping CIDRs. The peering request is rejected outright.
  3. n(n−1)/2 connections, each needing route table entries on both sides. Four VPCs is 6 connections; eight is 28. Every new VPC means editing every existing route table.

The local route is also not shared across a peering connection: you must add explicit routes on both sides, and security groups must permit the other VPC’s traffic. Security group referencing across a peering connection works, but only within a region and only if you enable it.

Peering remains the right answer for exactly two VPCs that will stay two — a shared-services VPC and one consumer, say. The moment a third arrives, do the arithmetic before adding a fourth peering connection.


Topic 4: Transit Gateway and Route Table Segmentation

A Transit Gateway is a router. Attachments plug into it; TGW route tables decide which attachments can reach which others. That segmentation is the actual product — without it, a TGW is just a more expensive full mesh.

The standard segmented design:

TGW route table: PROD
  attachments: prod-vpc-a, prod-vpc-b, shared-services, on-prem-vpn
  routes:      prod CIDRs, shared CIDRs, on-prem CIDRs
               (no route to dev — dev is unreachable, not merely denied)

TGW route table: DEV
  attachments: dev-vpc-a, dev-vpc-b, shared-services
  routes:      dev CIDRs, shared CIDRs
               (no on-prem, no prod)

TGW route table: SHARED
  attachments: shared-services
  routes:      everything — it must answer everyone

Each attachment is associated with one route table (which decides what it can reach) and can propagate its routes into several (which decides who can reach it). Getting that pair right is most of TGW operations, and misreading it is why “the route exists but traffic does not flow”.

Facts that shape designs:

  • Regional. Cross-region needs TGW peering, and TGW peering is not transitive either.
  • 50 Gbps per VPC attachment, aggregate across the attachment.
  • An attachment lands in specific subnets — one per AZ. If an AZ has no attachment subnet, instances in that AZ reach the TGW across AZs and you pay for it.
  • Appliance mode keeps a flow pinned to one AZ’s appliance for stateful inspection. Without it, asymmetric routing breaks any stateful firewall you insert.

Cost is the honest trade-off: per-attachment hourly plus per-GB processing, so a chatty pair of VPCs is cheaper peered. The usual mature design is a TGW for the general mesh, plus direct peering for the one or two very high-volume paths.


Topic 5: Hybrid — VPN and Direct Connect

Site-to-Site VPNDirect Connect
MediumIPsec over the internetDedicated physical circuit
SetupMinutesWeeks to months
Bandwidth~1.25 Gbps per tunnel1/10/100 Gbps
LatencyVariable — it is the internetConsistent
CostCheap hourly + dataPort fee + much cheaper egress
EncryptionAlwaysNone by default — run a VPN over it or use MACsec

Each VPN connection gives you two tunnels to two AWS endpoints for redundancy, and a surprising number of production VPNs run with only one configured. Configure both.

That last row of the table is the one that catches people: Direct Connect is a private circuit, not an encrypted one. Compliance regimes that require encryption in transit are not satisfied by “it is a dedicated line”.

The resilient hybrid pattern is DX as primary with a VPN as backup over the internet, both attached to the same Transit Gateway, with BGP preferring DX. When the circuit fails, BGP converges to the VPN and you take a bandwidth hit rather than an outage. Test the failover on purpose, on a schedule — an untested backup path is a hypothesis.


The same machinery that exposes AWS services privately can expose yours. Put a Network Load Balancer in front of your service, create a VPC endpoint service from it, and consumers create an interface endpoint in their own VPC.

Why this beats peering for a service:

  • No CIDR coordination. The consumer’s address space is irrelevant, and overlap does not matter.
  • One direction only. The consumer can reach your service; you cannot reach into their VPC. Peering is bidirectional by nature.
  • No route table changes on either side, and no blast radius growth as consumers are added.
  • You approve each consumer explicitly and can revoke one without affecting the others.

This is how SaaS vendors expose products inside your VPC, and it is the right internal pattern too — a platform team offering a service to twenty product teams should offer an endpoint service, not twenty peering connections.

Try it yourself: add the S3 gateway endpoint, then compare NAT gateway BytesOutToDestination in CloudWatch before and after. On a log-shipping workload the drop is dramatic and immediate, and it is the easiest cost win in this entire module.

Common mistake: creating an interface endpoint with private DNS enabled while a Route 53 private hosted zone for the same service name already exists. Both try to answer, one wins non-deterministically per resolver, and you get intermittent failures that correlate with nothing. Pick one mechanism per service name and delete the other.