AWS NAT Gateway Deep Dive - How It Works, Public vs Private NAT Gateway, Setup Step by Step, Limits (55,000 Connections, 100 Gbps), CloudWatch Metrics, Cost Optimisation, NAT Instance Comparison and Troubleshooting (AWS Part-14)


We created a NAT Gateway in Part-5 in one step and moved on. It deserves its own part, because the NAT Gateway is the component every private subnet depends on, it is the first surprise on most AWS bills, and the three most common "my instance cannot reach the internet" tickets all end at it.

This Part-14 goes under the hood - what NAT actually does to a packet, the public and private connectivity types, the exact limits, the metrics that tell you it is about to fail, what it costs and how to pay less, when a NAT instance still makes sense, how IPv6 changes the picture, and the complete troubleshooting list from the AWS docs. Numbers are from the current documentation - several (bandwidth, IP addresses per gateway) have gone up since I recorded the video.

Table of Content

  1. What network address translation does to a packet
  2. Public vs private NAT Gateway
  3. Set it up step by step - NAT Gateway, Elastic IP, route table
  4. High availability - one NAT Gateway per Availability Zone
  5. The hard numbers - bandwidth, connections, IP addresses, timeouts
  6. CloudWatch metrics to alarm on
  7. What a NAT Gateway costs
  8. Five ways to pay less
  9. NAT Gateway vs NAT instance
  10. IPv6 - NAT64, DNS64 and the egress-only Internet Gateway
  11. NAT Gateway with VPC peering and Transit Gateway
  12. The AWS CLI equivalents
  13. Troubleshooting - the official list, explained
  14. Conclusion



1. What network address translation does to a packet

An instance in a private subnet has only a private IP, say 10.0.11.20. Private addresses are not routable on the internet - no server out there can send a reply to 10.0.11.20. NAT (network address translation, specifically port address translation) fixes that by rewriting the packet on the way out and remembering how to rewrite the reply on the way back -

  1. 10.0.11.20:40001 sends a packet to 1.2.3.4:443. The private route table says 0.0.0.0/0 → nat-..., so it lands on the NAT Gateway.
  2. The gateway replaces the source with its own Elastic IP and a free port - 52.59.10.10:1025 - and writes 10.0.11.20:40001 ⇄ 52.59.10.10:1025 ⇄ 1.2.3.4:443 into its translation table.
  3. The packet goes out through the public subnet's route (0.0.0.0/0 → igw-...) and the Internet Gateway. The destination sees 52.59.10.10.
  4. The reply to 52.59.10.10:1025 comes back, the gateway looks up the entry, rewrites the destination to 10.0.11.20:40001, and forwards it.

Two consequences that explain half of the behaviour in this post - every private instance behind the gateway appears to the world as the same Elastic IP (great for vendor allow-lists), and nothing on the internet can start a connection inwards, because there is no table entry for it (that is why a NAT Gateway does not answer ping and cannot be used for inbound traffic).

NAT Gateway deep dive - the translation table, limits, metrics, and public vs private connectivity


2. Public vs private NAT Gateway

From the NAT gateway basics -

Public NAT GatewayPrivate NAT Gateway
Purposeprivate subnets reach the internetprivate subnets reach other VPCs or on-premises through a Transit Gateway or virtual private gateway
Elastic IPrequirednone - it uses a private IP from its subnet
Must sit ina public subnet (route to an Internet Gateway)any subnet
Typical useapt update, calling SaaS APIs, pulling container imagesconnecting networks with overlapping CIDRs - the whole VPC hides behind one private address the other side can route to

Both are managed, redundant within their AZ, support TCP, UDP and ICMP, and cannot have a security group (control traffic with the instances' security groups and the subnet NACLs - the gateway uses ports 1024-65535).



3. Set it up step by step - NAT Gateway, Elastic IP, route table

The full VPC build is in Part-5; the NAT part, with the current console -

  1. VPC → NAT gateways → Create NAT gateway. Name jhooq-nat-1a. Subnet - a public subnet (jhooq-public-1a). Connectivity type - Public. Elastic IP allocation ID - Allocate Elastic IP. Optionally Additional settings → Private IPv4 address to pin the gateway's private IP. Create NAT gateway.
  2. Wait for state Available (a minute or two). A Failed gateway shows the reason under State message and is auto-deleted after about an hour.
  3. Route tables → the private route table → Edit routes - 0.0.0.0/0 → NAT Gateway → jhooq-nat-1a. Make sure the private subnets are associated with that table.
  4. Confirm the public subnet's route table has 0.0.0.0/0 → igw-... - without it the gateway has no way out.
  5. Test from a private instance (Part-5, Step 7) - curl https://checkip.amazonaws.com prints the gateway's Elastic IP.

Give the EIP a Name tag as well, and write the address down - it is the one you hand to vendors who allow-list your traffic, and it stays yours until you release it.


4. High availability - one NAT Gateway per Availability Zone

A NAT Gateway is zonal - redundant inside its AZ, but if that AZ has an outage, the gateway is gone, and every private subnet routing through it loses internet access, including subnets in the healthy AZs. The docs are explicit: create a NAT gateway in each Availability Zone, and configure your routing to ensure that resources use the NAT gateway in the same Availability Zone. Concretely -

  1. jhooq-nat-1a in jhooq-public-1a, jhooq-nat-1b in jhooq-public-1b.
  2. Two private route tables - jhooq-private-rt-1a with 0.0.0.0/0 → jhooq-nat-1a, associated with jhooq-private-1a; jhooq-private-rt-1b with 0.0.0.0/0 → jhooq-nat-1b, associated with jhooq-private-1b.

This is also what the VPC and more wizard builds when you pick NAT gateways: 1 per AZ, and it removes the cross-AZ data charge you pay when traffic from 1b goes through a gateway in 1a. The price is one more gateway-hour; for production it is not optional.


5. The hard numbers - bandwidth, connections, IP addresses, timeouts

From the basics page, as of this update -

LimitValue
Bandwidth5 Gbps, scales automatically to 100 Gbps
Packets per second1 million, scales to 10 million - beyond that, packets are dropped
Simultaneous connections55,000 per IPv4 address per unique destination (destination IP + port + protocol)
IPv4 addresses per gatewayup to 8 (1 primary + 7 secondary) → up to 440,000 connections per destination; 2 Elastic IPs per public gateway by default, adjustable via Service Quotas
Idle timeout350 seconds - an idle connection is dropped and the next packet gets an RST
MTU8,500 bytes; keep instances at 1,500 for internet traffic; no IP fragmentation for TCP/ICMP
ProtocolsTCP, UDP, ICMP - no IPsec (use NAT-Traversal)
Per AZ5 NAT gateways per AZ by default
Security groupsnone - NACLs on the subnet, security groups on the instances

The 55,000 figure is the one that bites high-volume workloads - many instances all talking to the same destination (one API, one database endpoint) share the port range. The signal is the ErrorPortAllocation metric; the fix is secondary IPs or more gateways.



6. CloudWatch metrics to alarm on

NAT Gateways publish to CloudWatch under the AWS/NATGateway namespace at one-minute intervals, free. The ones worth an alarm -

  1. ErrorPortAllocation - the gateway could not allocate a source port. Anything above 0 means you are hitting the 55,000-per-destination limit. Alarm immediately; add secondary IPv4 addresses (Actions → Edit secondary IP address associations) or split subnets across more gateways.
  2. PacketsDropCount - packets dropped by the gateway; above 0 sustained means the gateway is overloaded or something unsupported (fragments, IPsec) is being sent.
  3. ActiveConnectionCount - trend it against 55,000 × number of IPs.
  4. IdleTimeoutCount - connections closed by the 350-second idle timeout; a rising count with application errors means you need TCP keepalives under 350 s.
  5. BytesOutToDestination / BytesInFromDestination - the GB that become your bill. A sudden step up is a backup job, a container pull loop or a compromised instance.
  6. ConnectionAttemptCount vs ConnectionEstablishedCount - a growing gap means destinations are refusing or unreachable.

To find out which private instance is generating the traffic, enable VPC Flow Logs on the NAT Gateway's network interface (EC2 → Network interfaces → description contains the nat- ID) and query by source address in CloudWatch Logs Insights.


7. What a NAT Gateway costs

From the pricing page - two meters, plus the IP -

  1. Hourly - charged for every hour the gateway exists, idle or not. About $0.045 per hour in us-east-1 ≈ $33 a month per gateway.
  2. Data processing - about $0.045 per GB that passes through, in either direction. 1 TB a month ≈ $46.
  3. Public IPv4 address - the Elastic IP, $0.005 per hour ≈ $3.60 a month.
  4. Plus the normal data transfer out charges to the internet, and cross-AZ charges if instances use a gateway in another AZ.

So two gateways (one per AZ) pushing 2 TB a month ≈ $66 + $92 + $7 ≈ $165 a month. For a startup's whole network that is often the single biggest line item - which is why the next section exists.


8. Five ways to pay less

The docs themselves list the first two; the rest are standard practice -

  1. Gateway endpoints for S3 and DynamoDB - free. Traffic to those services from private subnets goes through the endpoint instead of the NAT, so backups, logs, static assets and container image layers from ECR (stored in S3) stop costing $0.045 per GB. VPC → Endpoints → Create endpoint → com.amazonaws..s3 (Gateway) → tick the private route tables. If you do nothing else, do this.
  2. One gateway per AZ with per-AZ route tables - eliminates cross-AZ data charges on top of the HA benefit.
  3. Interface endpoints (PrivateLink) for the AWS services you talk to most - ECR API, CloudWatch Logs, SSM, Secrets Manager, STS. They cost per hour per AZ plus a small per-GB fee, which is still cheaper than NAT for heavy traffic, and they keep that traffic off the internet entirely. Compare with the PrivateLink pricing; the VPC endpoint parts (19 and 20) of this series go deep.
  4. Centralised egress - with a Transit Gateway, one set of NAT Gateways in a shared VPC serves all spoke VPCs; you pay TGW data processing but save N-1 gateways' hourly charges and get one place to inspect outbound traffic.
  5. IPv6 - IPv6-enabled workloads can reach IPv6 destinations through an egress-only Internet Gateway, which is free (section 10).

And the obvious one - delete lab gateways and set a billing alarm.



9. NAT Gateway vs NAT instance

Before the managed gateway existed (2015) you ran a NAT instance - an EC2 instance with IP forwarding and source/destination check disabled. AWS still documents it (NAT instances, comparison) and marks it as something you manage yourself -

NAT GatewayNAT instance
Availabilitymanaged, redundant in the AZyou build the failover (scripts, ASG)
Bandwidthup to 100 Gbpsdepends on the instance type
Maintenancenonepatching, monitoring, sizing
Security groupsnot supportedyes - you can filter on the NAT itself
Port forwarding / bastion / IPsecnoyes - it is a Linux box
Fragmented packetsnot supportedsupported
Cost~$33 per month + $0.045 per GBt4g.nano ≈ $3 per month, no per-GB processing fee (data transfer still applies)

For a dev or sandbox account where traffic is light and HA does not matter, a tiny NAT instance (the open-source fck-nat AMI is the popular packaged version) can cut the NAT bill by 90%. For production, the gateway's zero maintenance and scaling win. There is also the middle option for cheap labs - skip NAT entirely and give the single instance a public IP.


10. IPv6 - NAT64, DNS64 and the egress-only Internet Gateway

IPv6 addresses are globally routable, so an IPv6 subnet does not need NAT to reach the internet - but you still want "outbound only". That is the egress-only Internet Gateway: stateful like a NAT Gateway (replies come back, nothing can connect in), but free, and the route is ::/0 → eigw-... in the private route table.

The remaining problem is IPv6-only subnets that must reach IPv4-only services. That is what the NAT Gateway's NAT64 with DNS64 does - enable DNS64 on the subnet so that the Route 53 Resolver synthesises an IPv6 address (64:ff9b::/96 prefix) for IPv4-only hosts, route 64:ff9b::/96 → nat-..., and the NAT Gateway translates IPv6 to IPv4 on the way out. Same gateway, same pricing. Dual-stack subnets simply use the IGW for IPv6 and the NAT for IPv4.


11. NAT Gateway with VPC peering and Transit Gateway

Two rules from the basics page that explain mysterious timeouts -

  1. Over VPC peering - Client → NAT → Peering → Destination is supported (and return traffic finds its way back to the gateway), but Client → Peering → NAT → Internet is not - a peered VPC cannot use your NAT Gateway to reach the internet. That is the no edge-to-edge rule from Part-12.
  2. Over a Transit Gateway it is supported - which is exactly how centralised egress works: spoke VPCs route 0.0.0.0/0 to the TGW, the TGW routes it to the egress VPC attachment, and the egress VPC's NAT Gateways send it out. The return path needs routes back to the spoke CIDRs in the egress VPC's public route table. (A virtual private gateway - the older VPN endpoint - cannot do this; only TGW can.)


12. The AWS CLI equivalents

 1# public NAT gateway with a new EIP
 2EIP=$(aws ec2 allocate-address --domain vpc --query AllocationId --output text)
 3NAT=$(aws ec2 create-nat-gateway --subnet-id subnet-public1a --allocation-id "$EIP" \
 4  --connectivity-type public \
 5  --tag-specifications 'ResourceType=natgateway,Tags=[{Key=Name,Value=jhooq-nat-1a}]' \
 6  --query NatGateway.NatGatewayId --output text)
 7aws ec2 wait nat-gateway-available --nat-gateway-ids "$NAT"
 8
 9# private route table → NAT
10aws ec2 create-route --route-table-id rtb-private1a --destination-cidr-block 0.0.0.0/0 --nat-gateway-id "$NAT"
11
12# add a secondary EIP for more ports (quota: 2 EIPs per NAT by default)
13EIP2=$(aws ec2 allocate-address --domain vpc --query AllocationId --output text)
14aws ec2 associate-nat-gateway-address --nat-gateway-id "$NAT" --allocation-ids "$EIP2"
15
16# private NAT gateway (no EIP)
17aws ec2 create-nat-gateway --subnet-id subnet-private1a --connectivity-type private
18
19# metrics
20aws cloudwatch get-metric-statistics --namespace AWS/NATGateway --metric-name ErrorPortAllocation \
21  --dimensions Name=NatGatewayId,Value="$NAT" --statistics Sum --period 300 \
22  --start-time "$(date -u -d '1 hour ago' +%FT%TZ)" --end-time "$(date -u +%FT%TZ)"
23
24# delete, then release the EIP (otherwise it keeps billing)
25aws ec2 delete-nat-gateway --nat-gateway-id "$NAT"
26aws ec2 wait nat-gateway-deleted --nat-gateway-ids "$NAT"
27aws ec2 release-address --allocation-id "$EIP"

Terraform: aws_eip (domain = "vpc"), aws_nat_gateway (subnet_id, allocation_id, connectivity_type, secondary_allocation_ids), aws_route. The community terraform-aws-modules/vpc module has enable_nat_gateway, single_nat_gateway and one_nat_gateway_per_az flags that implement exactly the trade-offs in this post.


13. Troubleshooting - the official list, explained

The AWS troubleshooting page, with the one-line cause and fix for each -

1. NAT gateway creation fails (Failed state) - Read State message. Subnet has insufficient free addresses → free IPs or use another subnet. Network vpc-... has no internet gateway attached → attach an IGW first. Elastic IP eipalloc-... is already associated → disassociate it or allocate a new one. Failed gateways are deleted automatically after about an hour.

2. Performing this operation would exceed the limit of 5 NAT gateways - Per-AZ quota. Deleting gateways still count until state Deleted; request an increase or use another AZ.

3. The maximum number of addresses has been reached - Elastic IP quota (5 per region). Release unused ones (how to release an Elastic IP) or request an increase.

4. NotAvailableInZone - A constrained AZ; create the gateway in another AZ and route to it.

5. The gateway disappeared - It failed and was auto-deleted after an hour. Create it again and read the state message this time.

6. The gateway does not respond to ping - By design. Test from a private instance instead.

7. Instances cannot access the internet - In order: gateway state Available? gateway in a public subnet whose table routes 0.0.0.0/0 → igw? private subnet's table routes 0.0.0.0/0 → nat and is associated with the subnet? no more specific route sending traffic elsewhere? instance security group allows outbound (and ICMP for ping)? NACLs on both subnets allow outbound and the ephemeral return ports 1024-65535? protocol is TCP/UDP/ICMP? Flow logs show you which layer drops it.

8. Some TCP connections to one destination fail - The destination sends fragmented packets (unsupported - use a NAT instance for that destination) or has tcp_tw_recycle enabled (ask them to disable it).

9. traceroute does not show the NAT's private IP - A more specific route sends that traffic through another gateway (IGW, VPN). Check the route table for the instance's subnet.

10. Connection drops after 350 seconds - The idle timeout. Enable TCP keepalive on the instance with an interval under 350 s (net.ipv4.tcp_keepalive_time = 300), or send traffic.

11. IPsec cannot be established - Not supported through NAT; use NAT-T (UDP 4500) encapsulation.

12. Cannot initiate more connections / ErrorPortAllocation - The 55,000-per-destination limit. Add secondary IPv4 addresses (each gives another 55,000), create gateways per AZ or per subnet, close idle connections (IdleTimeoutCount), or limit client connection counts.


14. Conclusion

To summarise Part-14 -

  1. A NAT Gateway rewrites the source of outbound packets to its Elastic IP and keeps a translation table for the replies - so private instances share one public address, and nothing can connect in.
  2. Public gateways (EIP, public subnet) reach the internet; private gateways bridge overlapping networks through a Transit Gateway or virtual private gateway.
  3. It is zonal - run one per AZ with per-AZ route tables for availability and to avoid cross-AZ charges.
  4. Know the numbers - 5 to 100 Gbps, 55,000 connections per IP per destination (up to 8 IPs), 350 s idle timeout, no IPsec, no fragmentation - and alarm on ErrorPortAllocation and PacketsDropCount.
  5. It costs per hour and per GB - so add the free S3 and DynamoDB gateway endpoints, consider interface endpoints and centralised egress, use egress-only Internet Gateways for IPv6, and a NAT instance only for cheap labs.

The official references are the NAT gateway guide, NAT gateway basics, pricing and troubleshooting. The endpoints that replace much of the NAT traffic are the subject of the VPC endpoint and PrivateLink parts of this series, and the full VPC it lives in is Part-5.


AWS step by step series -

  1. Part-1 : AWS IAM user - create a user, group, policy, access keys and MFA
  2. Part-2 : AWS Organizations - multi-account setup, OUs and SCPs
  3. Part-3 : AWS assume IAM role - trust policy, switch role in console and CLI
  4. Part-4 : How to launch an EC2 instance - key pair, security group, SSH
  5. Part-5 : AWS VPC - public and private subnets, Internet Gateway, NAT Gateway, route tables
  6. Part-8 : EC2 launch template - versions, default version, source template, SSM parameter AMI
  7. Part-10 : EC2 Auto Scaling - launch template, Auto Scaling group, target tracking, ALB
  8. Part-11 : AWS WAF - web ACL, managed rules, rate limiting, geo blocking
  9. Part-12 : AWS VPC Peering - connect two VPCs, routes, security groups, DNS
  10. Part-13 : AWS Transit Gateway - hub-and-spoke for many VPCs and on-premises
  11. Part-14 : AWS NAT Gateway deep dive - public vs private, limits, cost, troubleshooting
  12. Part-15 : Amazon Route 53 - hosted zones, records, alias, routing policies, health checks
  13. Part-16 : AWS security groups - inbound and outbound rules, stateful, referencing, quotas
  14. Part-16 : AWS Certificate Manager - free TLS certificates for ALB, CloudFront and API Gateway
  15. Part-17 : AWS Lambda - function URLs, environment variables and layers
  16. Part-18 : Network Load Balancer - setup, and ALB vs NLB
  17. Part-19 : VPC endpoints - gateway and interface endpoints (PrivateLink) instead of NAT
  18. Part-20 : AWS PrivateLink - publish your own service with an endpoint service and NLB
  19. Part-20 : Amazon EBS volumes - types, attach, mount, resize, snapshots, encryption
  20. Part-21 : VPC Flow Logs - CloudWatch Logs, S3, record format, Logs Insights, Athena
  21. Part-21 : EC2 Spot Instances - pricing, interruptions, mixed instances groups
  22. Part-24 : AWS Control Tower - landing zone, controls, Account Factory, Identity Center

Networking fundamentals -

  1. What is a VPC and a subnet? AWS networking in five minutes
  2. What is CIDR? Calculate IP ranges for VPCs and subnets
  3. What is NAT? Static NAT, dynamic NAT and PAT explained

More AWS guides -

  1. What is AWS CloudFormation? Templates, stacks, change sets, drift, StackSets
  2. Learn AWS S3 - the complete course
  3. AWS API Gateway - REST API with Lambda, authorizers, Terraform
  4. AWS Advanced Networking Specialty (ANS-C01) - course companion
  5. AWS ECS and Fargate - how to deploy a Docker container
  6. AWS S3 - how to host a static website
  7. Terraform create EC2 instance on AWS
  8. Terraform AWS IAM - users, roles and policies
  9. Terraform and AWS multi-account setup
  10. Terraform - setting up an ALB and SSL

Posts in this series