AWS gives you two packet filters that look similar in the console and behave completely differently. Choosing wrongly does not usually cause an outage on the day you configure it — it causes one three months later when return traffic starts using a port nobody thought about.
Topic 1: Stateful vs Stateless
A security group is stateful. Allow inbound TCP/443 and the reply leaves automatically, from whatever ephemeral port the kernel chose, with no outbound rule required. Allow an outbound connection and the response comes back. Connection tracking does the work.
A network ACL is stateless. Each packet is evaluated on its own, and the reply is a separate packet in the other direction. Allow inbound 443 without allowing outbound 1024–65535 and every connection hangs — the request arrives, the application answers, and the answer is dropped on the way out.
That is why NACL rules always come in pairs, and why the ephemeral port range keeps appearing:
| Client OS | Ephemeral range |
|---|---|
| Linux kernels | 32768–60999 |
| Windows (modern) | 49152–65535 |
| NAT gateway | 1024–65535 |
| ELB | 1024–65535 |
Since you rarely control every client, 1024–65535 is what most NACLs end up allowing — which is a good illustration of how little a NACL is really filtering once it works.
Topic 2: The Full Comparison
| Security group | Network ACL | |
|---|---|---|
| Attaches to | An ENI (instance, ALB node, RDS, endpoint) | A subnet |
| State | Stateful | Stateless |
| Rules | Allow only | Allow and Deny |
| Evaluation | All rules, most permissive wins | Numbered, lowest match wins and stops |
| Default (new, custom) | Deny all in, allow all out | Deny all both ways |
Default (the default ones) | Allows itself as source | Allows everything |
| Can reference | Another security group, a prefix list | CIDR blocks only |
| Applies to | Only ENIs it is attached to | Every ENI in the subnet, no exceptions |
| Quota | 60 rules in + 60 out; 5 SGs per ENI | 20 rules per direction (40 hard max) |
The ordering difference matters when debugging. A security group evaluates every rule and allows the traffic if any of them permits it — there is no order and no way to write an exception. A NACL walks rules in numeric order, applies the first match, and stops: rule 100 allowing everything makes rule 200 denying one address dead code.
A NACL has no exceptions. It applies to every packet crossing the subnet boundary, including your own monitoring, your own SSM agent, and the reply to the request you are debugging with. This is exactly why a NACL that is “temporarily tightened” during an incident so often makes the incident worse.
Topic 3: Reference Groups, Not Addresses
The most useful security group feature is that the source of a rule can be another security group:
sg-alb inbound 443 from 0.0.0.0/0
sg-app inbound 8080 from sg-alb ← not a CIDR
sg-db inbound 5432 from sg-app ← not a CIDR
What this buys you:
- The rule keeps working when subnets change, when instances are replaced, when you add an AZ, and when the ASG scales. There is no address list to maintain.
- It expresses intent — “the app tier may reach the database” — which is what a reviewer needs to see, rather than a CIDR that requires cross-referencing to understand.
- It is tighter than a CIDR: another instance in the same subnet, without
sg-app, is denied. A CIDR-based rule would have let it through.
Two details that trip people up:
- A self-referencing rule (
sg-appallowingsg-app) is how you let cluster members talk to each other — Kafka, Elasticsearch, etcd. It is not automatic, except on thedefaultsecurity group. - Referencing works across peered VPCs only if you configure it, and never across regions. Across accounts the reference is written
111122223333/sg-0abc.
Managed prefix lists are the other maintainable pattern: define corp-offices once with your office CIDRs and reference it in every rule. Change an office IP in one place. AWS also publishes prefix lists for S3 and DynamoDB, which is what gateway endpoints put in your route table.
Topic 4: The Rules You Should Almost Never Write
✗ 0.0.0.0/0 : 22 SSH open to the internet.
Use SSM Session Manager. If you truly need SSH,
scope it to a prefix list of known addresses.
✗ 0.0.0.0/0 : 3389 RDP open to the internet. Same answer.
✗ 0.0.0.0/0 : 3306 A database exposed to the internet. There is no
0.0.0.0/0 : 5432 version of this that is correct.
✗ 0.0.0.0/0 : ALL "We'll tighten it later." Nobody has.
✗ 10.0.0.0/8 : ALL Flat internal trust — one compromised instance
reaches every other one.
Find them before an auditor does:
aws ec2 describe-security-groups --query \
"SecurityGroups[?IpPermissions[?IpRanges[?CidrIp=='0.0.0.0/0']]].[GroupId,GroupName]" \
--output table
Then check which of them are actually attached to something — an over-permissive group on nothing is a cleanup task, one on a production ENI is an incident:
aws ec2 describe-network-interfaces \
--filters Name=group-id,Values=sg-0abc \
--query 'NetworkInterfaces[].[NetworkInterfaceId,Description,PrivateIpAddress]' --output table
On outbound rules: the default is allow-all, and restricting it is a genuine security improvement — it is the difference between a compromised instance that can phone home and one that cannot. It is also work: you have to enumerate what the workload actually needs, including the AWS APIs it calls, and interface endpoints make that list much shorter and much more stable. Restrict egress on your data tier first, where the list is short and the value is highest.
Topic 5: When a NACL Is Actually the Right Tool
Given the ephemeral-port awkwardness, NACLs earn their place in a small number of cases:
- Blocking a specific hostile CIDR at the subnet edge, which a security group cannot express because it has no Deny.
- A regulatory requirement for a subnet-level control that exists independently of instance configuration.
- A hard isolation guarantee for a data subnet — a NACL that permits only the app-tier CIDRs is enforced regardless of any security group mistake made later.
Otherwise, leave the default NACL wide open and do the work in security groups. A common and defensible production posture is: default NACLs untouched, security groups doing all real filtering, and one deliberately tightened NACL on the data subnets.
Where else filtering can happen, so you know what you are looking at when a packet disappears:
internet
↓
[ NACL ] subnet boundary, stateless, both directions
↓
[ security group ] the ENI, stateful
↓
[ host firewall ] iptables/nftables/ufw on the instance — AWS cannot see it
↓
process and it may simply not be listening
Three layers before the application, and a fourth if the application binds to 127.0.0.1 instead of 0.0.0.0. Checking ss -ltnp early saves a surprising amount of time.
Topic 6: Proving It Rather Than Believing It
VPC flow logs record what the security groups and NACLs did, which is the only evidence that settles an argument:
2 111122223333 eni-0abc 10.20.64.5 10.20.128.9 43512 5432 6 12 3400 1755... 1755... ACCEPT OK
2 111122223333 eni-0abc 10.20.64.5 10.20.128.9 43512 5432 6 1 40 1755... 1755... REJECT OK
A REJECT on the inbound direction points at the destination’s security group or the NACL. An ACCEPT outbound with no matching return means the reply was dropped — which on a subnet with a NACL is almost always the missing ephemeral range.
For the fastest possible answer, Reachability Analyzer names the offending component directly rather than making you infer it. Between the two, you should never have to guess which layer dropped a packet.
Try it yourself: allow inbound 443 on a NACL and nothing outbound. Curl the instance. Watch the request arrive in the flow log and the response never leave. Then add outbound 1024–65535 and watch it work. Doing this once makes the stateless/stateful distinction permanent.
Common mistake: using CIDR-based security group rules that mirror the subnet layout, then changing the subnet layout. The rules still parse, still apply, and now permit a different set of machines than they used to — a security regression with no error message and no deployment. Reference security groups and the problem cannot occur.