There is no “public subnet” checkbox. A subnet is public because its route table sends 0.0.0.0/0 to an internet gateway, and for no other reason. Internalise that and half of VPC troubleshooting becomes reading two tables.
Topic 1: How the VPC Router Decides
Every subnet is associated with exactly one route table. If you do not associate one, it uses the VPC’s main route table — which is a common source of accidental exposure and accidental isolation in equal measure.
Every route table starts with a local route for the VPC CIDR that cannot be removed or overridden. That is why every subnet in a VPC can always reach every other subnet, regardless of what else you configure. Isolation between subnets is the job of security groups and NACLs, never of routing.
Longest prefix match wins. The router picks the most specific matching route, not the first:
Destination Target Chosen for...
10.20.0.0/16 local anything inside the VPC (always)
10.20.5.0/24 vpce-0abc overrides the local route for that /24
0.0.0.0/0 igw-0abc everything else
When two routes are equally specific, static routes beat propagated ones, and among dynamic sources the order is Direct Connect, then VPN. This ordering matters the moment you have both a VPN and Direct Connect to the same prefix — the traffic silently prefers DX, which is usually what you want and always worth knowing.
Topic 2: Internet Gateway vs NAT Gateway
The internet gateway is a VPC-level, horizontally scaled, fully redundant component — one per VPC, no bandwidth limit, no cost. It does two things: routes traffic to the internet, and performs the 1:1 NAT between an instance’s private IP and its public or Elastic IP.
That second part surprises people: the instance never sees its own public IP. ip addr shows the private address only. The translation happens at the gateway, which is why software that binds to or advertises its own IP needs to be told the public one explicitly, and why the instance metadata service exposes public-ipv4 as a separate field.
For an instance to reach the internet through an IGW, all four of these must be true:
- An internet gateway is attached to the VPC.
- The subnet’s route table has
0.0.0.0/0 → igw-xxx. - The instance has a public or Elastic IP.
- Security groups and NACLs permit the traffic in both directions.
Missing #3 is the classic case: the routing is perfect and the instance is silent, because there is no public address for return traffic to come back to.
The NAT gateway solves the opposite problem: private instances need outbound access — package updates, API calls, container pulls — with no inbound exposure. It performs many-to-one NAT, lives in a public subnet, and needs an Elastic IP. Private route tables point 0.0.0.0/0 at it.
Properties that decide your architecture:
- It is zonal. One NAT gateway per AZ, with each AZ’s private route table pointing at its own. A single shared NAT gateway means an AZ failure removes outbound access for the other AZs, and also means every byte from those AZs is billed as cross-AZ traffic.
- It scales to 100 Gbps and 55,000 simultaneous connections per unique destination. Hitting that ceiling is rare and produces
ErrorPortAllocationin CloudWatch — a metric worth an alarm because the symptom otherwise looks like random connection failures. - It is not free. An hourly charge plus a per-GB data processing charge, in every AZ, forever. On busy clusters the NAT bill is routinely the largest single line item in the networking section, and almost all of it is traffic to S3, ECR and other AWS services that should not be going through NAT at all.
NAT instance — the old DIY approach — still exists and is almost never right: you own the patching, the failover, the source/destination check, and the bandwidth ceiling. The one legitimate case is a very low-traffic sandbox where a t4g.nano costs less than a NAT gateway’s hourly rate.
Topic 3: The Route Table Patterns Worth Memorising
PUBLIC SUBNET route table
10.20.0.0/16 local
0.0.0.0/0 igw-0abc ← this line is what makes it "public"
pl-6da54004 vpce-s3-0abc ← S3 gateway endpoint (prefix list)
PRIVATE APP SUBNET route table (one per AZ)
10.20.0.0/16 local
0.0.0.0/0 nat-0abc-in-same-az ← outbound only, same-AZ NAT
pl-6da54004 vpce-s3-0abc ← S3 traffic skips the NAT entirely
10.99.0.0/16 tgw-0abc ← on-prem via Transit Gateway
PRIVATE DATA SUBNET route table
10.20.0.0/16 local
(no default route at all) ← a database has nothing to say
to the internet
That last table is worth pausing on. A data subnet with no 0.0.0.0/0 route cannot be exfiltrated to the internet even if the instance is compromised and even if a security group is wrong. It is the cheapest containment control in AWS, and it costs nothing but the discipline of using interface endpoints for the AWS APIs the database actually needs.
Egress-only internet gateway is the IPv6 equivalent of a NAT gateway: stateful, outbound-only, and free. IPv6 addresses in AWS are all globally routable, so an EIGW is the only way to have IPv6 outbound without IPv6 inbound.
Topic 4: Bastion Hosts Are a Legacy Pattern
The traditional design puts a hardened jump box in a public subnet with port 22 open to your office CIDR, and SSH agent forwarding to reach private instances. It works. It also means an internet-facing SSH port, a host to patch, a key distribution problem, and an audit trail that stops at the bastion.
Use SSM Session Manager instead. The instance holds an SSM agent (present on current Amazon Linux and Ubuntu AMIs) which makes an outbound connection to the SSM service. There is no inbound port, no bastion, no key:
aws ssm start-session --target i-0abc123
# Port forwarding — reach a private RDS instance from your laptop
aws ssm start-session --target i-0abc123 \
--document-name AWS-StartPortForwardingSessionToRemoteHost \
--parameters '{"host":["db.internal"],"portNumber":["5432"],"localPortNumber":["5432"]}'
What you gain: every session is authorised by IAM, logged in CloudTrail, and optionally recorded to S3 or CloudWatch Logs keystroke by keystroke. What you must configure: the instance profile needs AmazonSSMManagedInstanceCore, and a fully private instance needs three interface endpoints — ssm, ssmmessages, ec2messages. That endpoint requirement is the single most common reason “Session Manager doesn’t work”, and the symptom is the instance simply not appearing in the managed-instance list.
For the rare case where you genuinely need SSH, EC2 Instance Connect Endpoint gives you SSH to a private instance over an AWS-managed tunnel — again without a bastion or an open inbound port.
Topic 5: Reading a Broken Path
The order that finds the fault fastest, because each step eliminates a whole layer:
1. Is the instance running and its status checks passing?
aws ec2 describe-instance-status --instance-ids i-0abc
2. Which route table is this subnet actually associated with?
aws ec2 describe-route-tables \
--filters Name=association.subnet-id,Values=subnet-0abc
3. Does that table have the route you assumed it had?
(read it — do not assume the main table is not in play)
4. Security group: is the port open, from the right source?
5. NACL: are BOTH directions allowed, including ephemeral ports?
6. Is the process actually listening?
ss -ltnp on the instance
Then stop guessing and ask AWS directly. Reachability Analyzer evaluates the configuration path — every route table, security group, NACL and gateway between two ENIs — and tells you the exact component that blocks it, without sending a packet:
aws ec2 create-network-insights-path \
--source i-0abc --destination i-0def --protocol tcp --destination-port 443
aws ec2 start-network-insights-analysis --network-insights-path-id nip-0abc
aws ec2 describe-network-insights-analyses --network-insights-analysis-ids nia-0abc \
--query 'NetworkInsightsAnalyses[].[NetworkPathFound,Explanations[].ExplanationCode]'
ExplanationCode values like NO_ROUTE_TO_DESTINATION or ENI_SG_RULES_MISMATCH name the layer immediately. It costs a fraction of a cent per analysis and routinely replaces an hour of console clicking. It analyses configuration, not liveness — so a passing analysis with a failing connection means the fault is in the operating system or the application, which is itself a valuable answer.
Topic 6: Cutting the NAT Bill Without Weakening Anything
Because NAT charges per gigabyte, the routing decisions above are also cost decisions:
- An S3 gateway endpoint is free and removes S3 traffic from the NAT path entirely. On any cluster that reads objects or ships logs, this pays for itself immediately and reduces a dependency at the same time.
- A DynamoDB gateway endpoint is the same deal.
- Interface endpoints for ECR (
ecr.api,ecr.dkr) plus the S3 gateway endpoint stop container image pulls from crossing NAT. Image layers come from S3, which is why the S3 endpoint is required and not optional here. - Keep the NAT gateway in the same AZ as the instances using it. Cross-AZ NAT traffic is billed twice — once by NAT, once as inter-AZ transfer.
Confirm the effect rather than assuming it. Flow logs show the NAT gateway’s ENI; after adding an endpoint, S3 traffic should vanish from it.
Try it yourself: on a private instance, curl https://checkip.amazonaws.com and note the address — it is the NAT gateway’s Elastic IP, not the instance’s. Then remove the 0.0.0.0/0 route and try again. The instance is still perfectly reachable from inside the VPC, because the local route was never involved.
Common mistake: putting a NAT gateway in a private subnet. It launches, it shows as available, and nothing works — because the NAT gateway itself needs a route to an internet gateway, which only a public subnet has. The error surfaces as instances timing out on every outbound connection, and the NAT gateway looks healthy the entire time.