AWS Private Networking · 2026-10-03 · VPC Endpoints / PrivateLink arc
AWS Private Networking for MCP Servers: VPC Endpoints, PrivateLink, Transit Gateway, and Egress Filtering
An MCP server running in a private subnet without VPC endpoints does something surprising: every call to Secrets Manager, Bedrock, ECR, or CloudWatch Logs exits the VPC through a NAT Gateway, crosses the public internet, and re-enters the AWS network at the public service endpoint — costing $0.045/GB in NAT processing fees and adding unnecessary exposure to a public network path for credentials and model invocations. The five pillars of AWS private networking for MCP servers — VPC endpoints for zero-NAT AWS API calls, endpoint policies for resource-scoped access control, PrivateLink for cross-account MCP service exposure, Transit Gateway for multi-VPC hub topology, and Network Firewall for egress filtering — each contains sharp edges that cause silent failures or expensive misconfigurations. The default VPC endpoint policy allows any principal to call any resource in the service through the endpoint — tightening it is not automatic. A Network Firewall deployed without fixing all three route tables causes asymmetric routing where traffic enters through NAT and exits through NFW, producing connectivity failures with no firewall log entries. Transit Gateway propagation routes are enabled by default across all VPC attachments, which means dev and prod VPCs can route to each other unless you create custom route tables to enforce isolation. PrivateLink service acceptance is automatic by default for same-account consumers unless acceptance-required is set — external consumers can connect without approval. This guide synthesizes all five topics into three structural patterns: zero-NAT private API calls using VPC endpoints and endpoint policies, multi-VPC infrastructure using Transit Gateway and PrivateLink, and egress security using Network Firewall domain allowlists.
TL;DR
- Zero-NAT private API calls: add Gateway endpoints for S3 and DynamoDB (free — route table entries only, no ENI, no hourly charge). For Secrets Manager, ECR, Bedrock, SQS, Lambda, and CloudWatch Logs, create Interface endpoints with
--private-dns-enabledin each private subnet — the standard AWS hostname resolves to the endpoint ENI's private IP inside the VPC, so existing SDK code requires zero changes. Attach a security group to each Interface endpoint allowing inbound TCP 443 only from the MCP server's security group. Attach an endpoint policy to the Secrets Manager endpoint restricting actions toGetSecretValue/DescribeSecretfor specific secret ARN patterns only — the default policy allows any caller to reach any secret in the account through your endpoint. For S3, add a bucket policy withaws:SourceVpcedenying access from any request that did not come through the VPC endpoint — network-enforced boundary on top of IAM. - Multi-VPC hub (Transit Gateway + PrivateLink): create one TGW per region, attach shared-services VPC (centralized VPC endpoints), production VPC, and data VPC — each with one attachment per AZ subnet. Disable default route propagation (
--no-default-route-table-propagation) and create custom route tables per spoke group: prod-rt advertises prod VPC CIDR + shared-services VPC CIDR; dev-rt advertises dev VPC CIDR + shared-services but NOT prod — this enforces dev-prod isolation without security groups. Use AWS RAM to share the TGW across accounts. For cross-account MCP service access, expose the MCP server behind an internal NLB, create a VPC Endpoint Service (--acceptance-required), and allow consumer account principals explicitly — consumers get a named Interface endpoint with private DNS, reaching only the MCP service, with no peering and no CIDR overlap constraint. - Egress filtering (Network Firewall): create NFW in a dedicated firewall subnet, not the private subnet. Create a domain list rule group allowlisting only the hostnames your MCP server needs to call (
.anthropic.com,.amazonaws.com,api.github.com). Set firewall policy default action toALERT,aws:drop_strictfor stateful traffic to block everything not explicitly allowed. Fix the three route tables in the correct order: (1) private subnet route table:0.0.0.0/0→ NFW endpoint; (2) firewall subnet route table:0.0.0.0/0→ internet gateway; (3) internet gateway edge route table: private subnet CIDR → NFW endpoint. Run withALERTaction for 24–48 hours before switching unmatched traffic toDROP— new third-party dependencies show up in alert logs before you block them.
Pattern 1 — Zero-NAT Private API Calls (VPC Endpoints + Endpoint Policies)
Gateway endpoints for S3 and DynamoDB
Gateway endpoints are a deceptive name — they do not provision a gateway device. They work by injecting prefix list routes into subnet route tables: traffic destined for S3 or DynamoDB IP ranges is routed to the endpoint (which is just AWS backbone routing), bypassing the NAT Gateway entirely. There is no ENI, no security group, no hourly charge, no private DNS entry. For MCP servers in private subnets that make any S3 or DynamoDB calls, adding Gateway endpoints is pure cost reduction with zero reliability risk.
# Create Gateway endpoint for S3 (free)
aws ec2 create-vpc-endpoint \
--vpc-endpoint-type Gateway \
--service-name com.amazonaws.us-east-1.s3 \
--vpc-id vpc-0abc123 \
--route-table-ids rtb-0private-a rtb-0private-b
# Create Gateway endpoint for DynamoDB (free)
aws ec2 create-vpc-endpoint \
--vpc-endpoint-type Gateway \
--service-name com.amazonaws.us-east-1.dynamodb \
--vpc-id vpc-0abc123 \
--route-table-ids rtb-0private-a rtb-0private-b
# Verify route table entries were injected
aws ec2 describe-route-tables \
--route-table-ids rtb-0private-a \
--query 'RouteTables[0].Routes[?starts_with(GatewayId, `vpce-`)]'
# Each entry: DestinationPrefixListId (pl-xxx) + GatewayId (vpce-xxx)
Only add Gateway endpoint routes to the private subnet route tables — not public subnets. Public subnets reach S3 through the internet gateway; adding a Gateway endpoint there is redundant. For multi-AZ deployments with two private route tables (one per AZ), both route tables must be in the --route-table-ids list or only one AZ gets the cost savings. The prefix list routes do not affect DNS resolution — s3.amazonaws.com still resolves to public S3 IPs; it is the routing that bypasses NAT when traffic matches the prefix list.
Interface endpoints for Secrets Manager, Bedrock, ECR, and CloudWatch Logs
Interface endpoints provision an ENI in each subnet specified. The ENI gets a private IP in the subnet's CIDR range. When private DNS is enabled (the default for most services, and the configuration you almost always want), AWS overrides the VPC's DNS resolution for the service hostname: secretsmanager.us-east-1.amazonaws.com resolves to the endpoint ENI's private IP inside the VPC rather than the public Secrets Manager IP. Your MCP server's existing AWS SDK calls route to the endpoint with zero code changes.
# Create Interface endpoint for Secrets Manager with private DNS
aws ec2 create-vpc-endpoint \
--vpc-endpoint-type Interface \
--service-name com.amazonaws.us-east-1.secretsmanager \
--vpc-id vpc-0abc123 \
--subnet-ids subnet-0private-a subnet-0private-b \
--security-group-ids sg-0mcp-endpoint \
--private-dns-enabled
# Create Interface endpoint for Amazon Bedrock Runtime (LLM invocations)
aws ec2 create-vpc-endpoint \
--vpc-endpoint-type Interface \
--service-name com.amazonaws.us-east-1.bedrock-runtime \
--vpc-id vpc-0abc123 \
--subnet-ids subnet-0private-a subnet-0private-b \
--security-group-ids sg-0mcp-endpoint \
--private-dns-enabled
# Create Interface endpoint for ECR (container image pulls)
# ECR requires TWO endpoints: ecr.api and ecr.dkr
aws ec2 create-vpc-endpoint \
--vpc-endpoint-type Interface \
--service-name com.amazonaws.us-east-1.ecr.api \
--vpc-id vpc-0abc123 \
--subnet-ids subnet-0private-a subnet-0private-b \
--security-group-ids sg-0mcp-endpoint \
--private-dns-enabled
aws ec2 create-vpc-endpoint \
--vpc-endpoint-type Interface \
--service-name com.amazonaws.us-east-1.ecr.dkr \
--vpc-id vpc-0abc123 \
--subnet-ids subnet-0private-a subnet-0private-b \
--security-group-ids sg-0mcp-endpoint \
--private-dns-enabled
# Create Interface endpoint for CloudWatch Logs
aws ec2 create-vpc-endpoint \
--vpc-endpoint-type Interface \
--service-name com.amazonaws.us-east-1.logs \
--vpc-id vpc-0abc123 \
--subnet-ids subnet-0private-a subnet-0private-b \
--security-group-ids sg-0mcp-endpoint \
--private-dns-enabled
The security group attached to the endpoint (sg-0mcp-endpoint) controls inbound traffic to the endpoint ENI. The minimal correct rule: allow inbound TCP 443 only from the MCP server's security group. The endpoint SG needs no outbound rules — response traffic is handled by the stateful tracking. ECR is a common surprise: pulling container images requires both ecr.api (for API operations like GetAuthorizationToken) and ecr.dkr (for the Docker protocol layer transfer). Creating only one results in partial failures — the auth token fetch works but the layer download hangs. Additionally, ECR image layers are stored in S3, so the Gateway S3 endpoint is required for image pulls to succeed without NAT.
Cost note: Interface endpoints are $0.01/hour per AZ per endpoint. A deployment in two AZs with four Interface endpoints (Secrets Manager, Bedrock Runtime, ECR API, ECR DKR) plus CloudWatch Logs costs $0.10/hour = $73/month. This is almost always less than the equivalent NAT Gateway data processing fees at any meaningful API call volume.
Endpoint policies as second-layer access control
The default endpoint policy is Allow * on all actions and all resources — completely permissive. This means any principal that has IAM permissions to call S3 or Secrets Manager through the endpoint can reach any bucket or secret in any account accessible from the VPC. An endpoint policy constrains what can be reached through the endpoint, independent of IAM. Both must allow a call — the intersection governs. See the endpoint policy reference for the full policy evaluation model and IAM-endpoint interaction diagram.
# Restrict Secrets Manager endpoint to specific secret ARN patterns
aws ec2 modify-vpc-endpoint \
--vpc-endpoint-id vpce-0secrets123 \
--policy-document '{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "AllowMCPServerSecrets",
"Effect": "Allow",
"Principal": "*",
"Action": [
"secretsmanager:GetSecretValue",
"secretsmanager:DescribeSecret"
],
"Resource": "arn:aws:secretsmanager:us-east-1:123456789012:secret:mcp-server-*"
}
]
}'
# Restrict S3 Gateway endpoint to specific buckets
aws ec2 modify-vpc-endpoint \
--vpc-endpoint-id vpce-0s3gateway123 \
--policy-document '{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "AllowMCPBuckets",
"Effect": "Allow",
"Principal": "*",
"Action": "s3:*",
"Resource": [
"arn:aws:s3:::mcp-server-data",
"arn:aws:s3:::mcp-server-data/*",
"arn:aws:s3:::mcp-assets",
"arn:aws:s3:::mcp-assets/*"
]
}
]
}'
# S3 bucket policy: deny access if request does not come via VPC endpoint
aws s3api put-bucket-policy \
--bucket mcp-server-data \
--policy '{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "DenyNonVPCEAccess",
"Effect": "Deny",
"Principal": "*",
"Action": "s3:*",
"Resource": [
"arn:aws:s3:::mcp-server-data",
"arn:aws:s3:::mcp-server-data/*"
],
"Condition": {
"StringNotEquals": {
"aws:SourceVpce": "vpce-0s3gateway123"
}
}
}
]
}'
The S3 bucket policy with aws:SourceVpce is the strongest enforcement mechanism: it makes the bucket unreachable from any path other than your specific VPC endpoint — not even from the console or CLI outside the VPC unless the caller is coming through that endpoint. This is the network-enforced boundary that IAM alone cannot provide, because IAM policies control identity but not network path. The endpoint policy restricts what resources are reachable through the endpoint; the bucket policy restricts which network paths can reach the bucket. Combined, both must be satisfied.
One critical operational gotcha: when using aws:SourceVpce in a bucket policy, IAM administrators, CI/CD pipelines, and any tooling running outside the VPC will get AccessDenied unless they're routing through the endpoint. Always test access patterns before applying DenyNonVPCEAccess in production. The typical pattern is to deploy the bucket policy with a Deny, audit the CloudTrail access logs for denied calls, fix each caller, then confirm the policy is correct.
Pattern 2 — Multi-VPC Architecture (Transit Gateway + PrivateLink)
Transit Gateway hub topology for MCP infrastructure
When MCP server infrastructure spans more than two VPCs — production ECS cluster, staging, shared-services (centralized VPC endpoints), data layer (RDS, DynamoDB, Elasticsearch) — VPC peering becomes the wrong tool. Peering is non-transitive: if production peers with shared-services and shared-services peers with data, production cannot reach data through shared-services without a direct peering connection. With five VPCs you need 10 peering connections. Peering also fails when VPCs have overlapping CIDR blocks, which happens in organizations that didn't plan CIDRs centrally. See Transit Gateway architecture patterns for the full route table design and RAM sharing.
# Create Transit Gateway (one per region)
TGW_ID=$(aws ec2 create-transit-gateway \
--description "MCP infrastructure hub" \
--options '{
"AmazonSideAsn": 64512,
"AutoAcceptSharedAttachments": "disable",
"DefaultRouteTableAssociation": "disable",
"DefaultRouteTablePropagation": "disable",
"VpnEcmpSupport": "enable",
"DnsSupport": "enable"
}' \
--query 'TransitGateway.TransitGatewayId' --output text)
# Attach shared-services VPC (centralized VPC endpoints)
aws ec2 create-transit-gateway-vpc-attachment \
--transit-gateway-id $TGW_ID \
--vpc-id vpc-0shared-services \
--subnet-ids subnet-0shared-a subnet-0shared-b \
--options "ApplianceModeSupport=disable,DnsSupport=enable,Ipv6Support=disable"
# Attach production VPC
aws ec2 create-transit-gateway-vpc-attachment \
--transit-gateway-id $TGW_ID \
--vpc-id vpc-0prod \
--subnet-ids subnet-0prod-a subnet-0prod-b
# Attach data VPC (RDS, ElastiCache)
aws ec2 create-transit-gateway-vpc-attachment \
--transit-gateway-id $TGW_ID \
--vpc-id vpc-0data \
--subnet-ids subnet-0data-a subnet-0data-b
The two critical flags are DefaultRouteTableAssociation: disable and DefaultRouteTablePropagation: disable. With defaults enabled, every attachment is associated with the default route table and propagates its CIDR to all other attachments — meaning every VPC can reach every other VPC from day one, including dev-to-prod paths that should be isolated. Disabling these forces explicit route table design.
# Create custom route tables for isolation
# "prod-rt": prod can reach shared-services and data, but not dev
PROD_RT=$(aws ec2 create-transit-gateway-route-table \
--transit-gateway-id $TGW_ID \
--tag-specifications 'ResourceType=transit-gateway-route-table,Tags=[{Key=Name,Value=prod-rt}]' \
--query 'TransitGatewayRouteTable.TransitGatewayRouteTableId' --output text)
# "shared-rt": shared-services can reach prod and data (hub — receives from all)
SHARED_RT=$(aws ec2 create-transit-gateway-route-table \
--transit-gateway-id $TGW_ID \
--tag-specifications 'ResourceType=transit-gateway-route-table,Tags=[{Key=Name,Value=shared-rt}]' \
--query 'TransitGatewayRouteTable.TransitGatewayRouteTableId' --output text)
# Associate attachments with their route tables
# Production VPC attachment → prod-rt
aws ec2 associate-transit-gateway-route-table \
--transit-gateway-route-table-id $PROD_RT \
--transit-gateway-attachment-id tgw-attach-0prod
# Enable propagation: prod-rt receives routes from shared-services and data
aws ec2 enable-transit-gateway-route-table-propagation \
--transit-gateway-route-table-id $PROD_RT \
--transit-gateway-attachment-id tgw-attach-0shared
aws ec2 enable-transit-gateway-route-table-propagation \
--transit-gateway-route-table-id $PROD_RT \
--transit-gateway-attachment-id tgw-attach-0data
# Dev attachments go to dev-rt which only propagates from shared-services
# dev-rt does NOT propagate from prod-attachment = dev cannot reach prod
Propagation populates the route table with routes learned from each attachment. Association determines which route table a VPC attachment uses for inbound routing. They are independent: you can have prod-rt propagate routes from shared-services and data (so prod can reach both) without propagating from dev (so prod cannot reach dev). The dev-rt does the reverse — dev can reach shared-services for VPC endpoints and monitoring, but prod routes are never propagated into dev-rt, so dev MCP instances cannot reach the production database.
Cross-account sharing uses AWS Resource Access Manager. The networking team owns the TGW in the central networking account and shares it with product team accounts. Product team accounts create VPC attachments to the shared TGW — they don't need their own TGW and don't need permissions to the TGW itself beyond creating attachments.
# Share TGW with the organization or specific accounts
aws ram create-resource-share \
--name mcp-tgw-share \
--resource-arns arn:aws:ec2:us-east-1:111111111111:transit-gateway/$TGW_ID \
--principals arn:aws:organizations::111111111111:organization/o-xxxx
# Or use specific account IDs: --principals 222222222222 333333333333
PrivateLink for cross-account MCP service exposure
PrivateLink solves a different problem than TGW: when you need consumers in other accounts to reach a specific MCP service endpoint without any network-level access to the provider VPC. TGW merges routing between VPCs; PrivateLink is unidirectional and service-scoped — the consumer gets access to one named service endpoint and nothing else in the provider VPC, regardless of security group rules. No CIDR conflicts, no route table merges, no lateral movement risk. See PrivateLink cross-account patterns for the full consumer-side setup and DNS options.
# Provider side: MCP server must be behind an internal NLB
NLB_ARN=$(aws elbv2 create-load-balancer \
--name mcp-internal-nlb \
--type network \
--scheme internal \
--subnets subnet-0private-a subnet-0private-b \
--query 'LoadBalancers[0].LoadBalancerArn' --output text)
TG_ARN=$(aws elbv2 create-target-group \
--name mcp-server-tg \
--protocol TCP \
--port 3000 \
--vpc-id vpc-0provider \
--target-type ip \
--health-check-protocol TCP \
--query 'TargetGroups[0].TargetGroupArn' --output text)
aws elbv2 create-listener \
--load-balancer-arn $NLB_ARN \
--protocol TCP \
--port 3000 \
--default-actions Type=forward,TargetGroupArn=$TG_ARN
# Create the VPC Endpoint Service
SVC_ID=$(aws ec2 create-vpc-endpoint-service-configuration \
--network-load-balancer-arns $NLB_ARN \
--acceptance-required \
--query 'ServiceConfiguration.ServiceId' --output text)
# Get the service name consumers will use
SERVICE_NAME=$(aws ec2 describe-vpc-endpoint-service-configurations \
--service-ids $SVC_ID \
--query 'ServiceConfigurations[0].ServiceName' --output text)
# e.g., com.amazonaws.vpce.us-east-1.vpce-svc-0abc123
# Allow specific consumer account principals
aws ec2 modify-vpc-endpoint-service-permissions \
--service-id $SVC_ID \
--add-allowed-principals arn:aws:iam::222222222222:root
The --acceptance-required flag is critical for external consumer access control. Without it, any principal in the allowed-principals list can create an endpoint and start sending traffic immediately — no provider notification, no approval step. With --acceptance-required, each new consumer endpoint request shows up in the console and via EventBridge, and the provider must explicitly accept it before traffic flows. For intra-organization use where consumers are known accounts, you can set --acceptance-required false to allow auto-accept for convenience.
# Consumer side: create Interface endpoint pointing at the service
aws ec2 create-vpc-endpoint \
--vpc-endpoint-type Interface \
--service-name $SERVICE_NAME \
--vpc-id vpc-0consumer \
--subnet-ids subnet-0consumer-a subnet-0consumer-b \
--security-group-ids sg-0consumer-mcp \
--private-dns-enabled false
# Note: private-dns-enabled false for custom service names —
# provider must own the domain to enable private DNS.
# Consumer gets a VPC-local DNS name: vpce-xxx.vpce-svc-xxx.us-east-1.vpce.amazonaws.com
# For branded DNS: provider associates a Route 53 private hosted zone
aws ec2 associate-vpc-with-hosted-zone \
--vpc-id vpc-0provider \
--hosted-zone-id Z0ABCDEF123456
# Consumer VPC also needs association for split-horizon resolution
PrivateLink does not route general VPC-to-VPC traffic — it is service-specific. The consumer endpoint resolves to a private IP in the consumer VPC (the endpoint ENI), and all traffic to the MCP service goes through that ENI to the provider NLB without touching the provider VPC's subnets or security groups. An attacker who compromises the consumer-side MCP client cannot pivot to other services in the provider VPC through PrivateLink — they can only reach the NLB target group listeners.
Pattern 3 — Egress Filtering (Network Firewall)
Why VPC endpoints alone are not enough for egress control
VPC endpoints keep intra-AWS traffic off the public internet, but they do not control outbound traffic to non-AWS destinations. An MCP server that calls external LLM APIs (api.anthropic.com, api.openai.com), third-party data providers, or webhooks must exit through a NAT Gateway or internet gateway. If that container is compromised — through a prompt injection, a vulnerable dependency, or a misconfigured IAM role — the attacker can exfiltrate data to arbitrary internet destinations using the same NAT path. AWS Network Firewall sits in the network path and inspects all egress traffic, blocking connections to everything not on an explicit allowlist. See Network Firewall configuration patterns for the full routing architecture and Suricata rule reference.
Network Firewall deployment and the three route tables
The hardest part of Network Firewall deployment is getting the route tables right. NFW does not operate as a transparent in-line device — it requires three distinct route table configurations that must all be consistent. Misconfiguring any one of them results in asymmetric routing (traffic goes one way through NFW, returns a different way), which TCP treats as a dropped connection.
# Step 1: Create NFW in a dedicated firewall subnet (separate from private subnets)
# Each AZ needs its own firewall subnet (e.g., 10.0.3.0/28 in AZ-a, 10.0.4.0/28 in AZ-b)
NFW_ID=$(aws network-firewall create-firewall \
--firewall-name mcp-egress-fw \
--firewall-policy-arn $POLICY_ARN \
--vpc-id vpc-0mcp-prod \
--subnet-mappings '[
{"SubnetId": "subnet-0fw-a"},
{"SubnetId": "subnet-0fw-b"}
]' \
--query 'Firewall.FirewallId' --output text)
# Get the endpoint IDs for each AZ (needed for route table entries)
aws network-firewall describe-firewall \
--firewall-name mcp-egress-fw \
--query 'FirewallStatus.SyncStates'
# Output: { "us-east-1a": { "Attachment": { "EndpointId": "vpce-0fw-endpoint-a" } }, ... }
# Step 2: Create the firewall policy with domain allowlist
# First, create a domain list stateful rule group
RULE_GROUP_ARN=$(aws network-firewall create-rule-group \
--rule-group-name mcp-egress-domains \
--type STATEFUL \
--capacity 100 \
--rule-group '{
"StatefulRulesAndCustomActions": {
"StatefulRules": [],
"RulesSourceList": {
"Targets": [
".anthropic.com",
".openai.com",
".amazonaws.com",
"api.github.com",
"registry.npmjs.org"
],
"TargetTypes": ["HTTP_HOST", "TLS_SNI"],
"GeneratedRulesType": "ALLOWLIST"
}
}
}' \
--query 'RuleGroupResponse.RuleGroupArn' --output text)
# Create firewall policy with default DROP for unmatched stateful traffic
aws network-firewall create-firewall-policy \
--firewall-policy-name mcp-egress-policy \
--firewall-policy '{
"StatelessDefaultActions": ["aws:forward_to_sfe"],
"StatelessFragmentDefaultActions": ["aws:forward_to_sfe"],
"StatefulDefaultActions": ["aws:alert_strict"],
"StatefulEngineOptions": {
"RuleOrder": "DEFAULT_ACTION_ORDER"
},
"StatefulRuleGroupReferences": [
{
"ResourceArn": "'"$RULE_GROUP_ARN"'"
}
]
}'
The aws:alert_strict default action alerts on all traffic not matched by a rule and then drops it. During the initial rollout, use aws:alert_established (alert-only, no drop) for 24–48 hours. Monitor the alert logs in CloudWatch to discover all outbound destinations the MCP server is actually calling before enabling the strict drop. Switching too quickly to drop mode breaks unanticipated dependencies — npm registry calls, OS package updates, third-party SDK telemetry endpoints.
# Step 3: Fix route tables — ALL THREE must be correct simultaneously
# Route table 1: private subnet (MCP server's subnet)
# Default route → NFW endpoint in same AZ
aws ec2 create-route \
--route-table-id rtb-0private-a \
--destination-cidr-block 0.0.0.0/0 \
--vpc-endpoint-id vpce-0fw-endpoint-a # AZ-a endpoint
# Route table 2: firewall subnet (where NFW is deployed)
# Default route → internet gateway
aws ec2 create-route \
--route-table-id rtb-0firewall-a \
--destination-cidr-block 0.0.0.0/0 \
--gateway-id igw-0main
# Route table 3: IGW edge route table (attached to the internet gateway)
# Return traffic for private subnet CIDR → NFW endpoint
aws ec2 create-route \
--route-table-id rtb-0igw-edge \
--destination-cidr-block 10.0.1.0/24 \ # private subnet AZ-a CIDR
--vpc-endpoint-id vpce-0fw-endpoint-a
# Attach the edge route table to the internet gateway
aws ec2 associate-route-table \
--route-table-id rtb-0igw-edge \
--gateway-id igw-0main
The asymmetric routing failure mode: if the IGW edge route table is missing or points to the wrong NFW endpoint (e.g., AZ-a returns traffic via AZ-b's endpoint), TCP connections fail because the SYN goes out via one NFW instance and the SYN-ACK arrives via a different one. Each NFW endpoint maintains its own connection table — cross-AZ return traffic lands on an NFW instance that has no record of the connection and drops it. The fix is strict AZ affinity: private subnet in AZ-a → NFW endpoint in AZ-a → IGW, and return traffic for AZ-a's CIDR → NFW endpoint in AZ-a.
Monitoring Network Firewall alerts before enabling drop mode
Network Firewall publishes alert logs to S3, CloudWatch Logs, or Kinesis Firehose. Alert log entries include the source IP, destination IP, destination domain (for TLS SNI / HTTP Host matches), protocol, port, and rule that generated the alert. During the 24–48 hour observation window before enabling DROP, query the alert logs to build the complete list of outbound destinations.
# Query NFW alert logs in CloudWatch Logs (if log destination is CWL)
aws logs start-query \
--log-group-name /aws/network-firewall/mcp-egress-fw/alert \
--start-time $(date -d '24 hours ago' +%s) \
--end-time $(date +%s) \
--query-string '
fields @timestamp, event.dest_ip, event.http.hostname, event.tls.sni, event.proto, event.dest_port
| filter event.event_type = "alert"
| stats count(*) as hit_count by event.tls.sni, event.http.hostname, event.dest_port
| sort hit_count desc
| limit 100
'
# Check query results
aws logs get-query-results --query-id $QUERY_ID
Common surprise destinations found during the observation window: AWS SDK retry telemetry (back to *.amazonaws.com — already on the allowlist), npm registry subdomains (registry.npmjs.org, cdn.npmjs.org), Docker Hub (registry-1.docker.io, auth.docker.io), and third-party observability agents (Datadog, Sentry, New Relic). Add each legitimate destination to the domain list rule group before switching to DROP mode. Document why each domain is needed — this list becomes the egress access control policy for your MCP server.
Failure modes reference
| Component | Symptom | Root cause | Fix |
|---|---|---|---|
| S3 Gateway endpoint | S3 calls still processed by NAT Gateway (high data transfer costs) | Gateway endpoint added to public subnet route table, not private subnet route table | Add --route-table-ids for private subnet route tables only |
| ECR Interface endpoint | Container pull fails with "no basic auth credentials" | Only ecr.dkr created; ecr.api endpoint missing so GetAuthorizationToken fails |
Create both ecr.api and ecr.dkr endpoints; also verify S3 Gateway endpoint for image layers |
| Interface endpoint private DNS | MCP server SDK calls go through NAT despite endpoint existing | Endpoint created without --private-dns-enabled; hostname still resolves to public IP |
Delete endpoint and recreate with --private-dns-enabled, or modify to enable private DNS |
| Endpoint security group | SDK calls to Secrets Manager time out with no error | Endpoint SG does not allow inbound TCP 443 from MCP server SG | Add inbound rule: TCP 443 from MCP server security group ID |
| Endpoint policy (Secrets Manager) | GetSecretValue returns AccessDenied despite correct IAM role |
Custom endpoint policy was set with wrong principal filter or missing resource ARN pattern | Check endpoint policy allows the calling IAM role ARN or use Principal: "*" with resource restrictions only |
| S3 bucket policy (SourceVpce) | CI/CD S3 sync or console access returns AccessDenied |
Deny condition applies to all callers not using the VPC endpoint, including external CI/CD | Add exception for CI/CD IAM role ARN using aws:PrincipalArn condition, or route CI/CD through the VPC |
| Transit Gateway route propagation | Dev VPC can reach production database | Default route table propagation enabled; all VPC CIDRs propagated to all attachments | Create custom route tables per spoke group; only propagate routes from permitted attachments |
| Transit Gateway attachment | Cross-VPC traffic fails despite TGW route table showing correct routes | VPC route table missing 0.0.0.0/0 or specific CIDR route pointing to TGW attachment |
Add route in each VPC's private subnet route table: remote CIDR → TGW attachment ID |
| PrivateLink acceptance | Consumer endpoint stuck in pending-acceptance state indefinitely |
Provider set --acceptance-required but no one is monitoring the acceptance queue |
Create EventBridge rule on com.amazonaws.us-east-1.ec2 / AWS API Call via CloudTrail / CreateVpcEndpoint to alert the provider team |
| PrivateLink NLB target health | Consumer can create endpoint but MCP requests time out | NLB target group health checks failing — MCP container not listening on expected health check port | Set --health-check-port to the MCP server's actual health endpoint port; verify SG allows NLB health check traffic |
| Network Firewall routing | All egress from private subnet times out after NFW deployment | Asymmetric routing: one of the three route tables (private, firewall, IGW edge) is missing or incorrect | Verify all three route tables; ensure IGW edge route table is attached to the internet gateway |
| Network Firewall cross-AZ | NFW connectivity works in one AZ but not the other | IGW edge route table has only one AZ's private subnet CIDR → NFW endpoint; AZ-b's return traffic routes to AZ-a's NFW endpoint | Add IGW edge route for each AZ's private subnet CIDR pointing to that AZ's NFW endpoint |
| Network Firewall domain list | HTTPS connections to allowlisted domain time out intermittently | Domain list uses HTTP_HOST match type only; TLS connections use SNI, not HTTP Host header |
Use both HTTP_HOST and TLS_SNI in TargetTypes for HTTPS destinations |
| Network Firewall alert-to-drop transition | MCP server breaks unexpectedly after enabling DROP mode | Undiscovered dependency (npm package CDN, observability telemetry, OS update mirror) not in domain allowlist | Extend alert-only observation window; query alert logs for all unique destinations before switching to DROP |
| VPC endpoint + Network Firewall conflict | S3 calls fail after NFW deployment despite S3 Gateway endpoint existing | Private subnet default route now points to NFW instead of NAT; Gateway endpoint route is more specific but applies only to S3 prefix lists — need to verify route priority | Gateway endpoint routes are prefix list routes (more specific than 0.0.0.0/0) and take precedence over the NFW default route — check the route table for the S3 prefix list entry; if missing, re-add the Gateway endpoint route table association |
Production checklists
VPC endpoints setup checklist
- S3 and DynamoDB Gateway endpoints added to all private subnet route tables where MCP servers run (not public subnets)
- Interface endpoints created for: Secrets Manager, Bedrock Runtime, ECR API, ECR DKR, CloudWatch Logs — each with
--private-dns-enabled - Each Interface endpoint has a security group allowing inbound TCP 443 from MCP server security group only
- Interface endpoints deployed in all private subnets (one ENI per AZ) to avoid single-AZ dependency
- Secrets Manager endpoint policy restricts
GetSecretValuetomcp-server-*ARN pattern — defaultAllow *replaced - S3 Gateway endpoint policy restricts
s3:*to specific bucket ARNs — not all S3 in the account - S3 bucket policy includes
DenyNonVPCEAccesscondition withaws:SourceVpce— tested that CI/CD and console access still work (route through VPC or principal exception) - NAT Gateway data processing costs in CloudWatch trending down after endpoint deployment
Transit Gateway checklist
- TGW created with
DefaultRouteTableAssociation: disableandDefaultRouteTablePropagation: disable - Custom route tables created per spoke group (prod-rt, dev-rt, data-rt, shared-rt)
- Propagation configured explicitly: prod-rt propagates from shared-services and data; dev-rt propagates from shared-services only — verified no prod CIDRs in dev-rt
- VPC route tables in each VPC updated with routes to remote VPC CIDRs via TGW attachment
- TGW shared via AWS RAM with all product team accounts — product team accounts can create attachments
- Verified cross-VPC connectivity with
aws ec2 create-flow-logsand a test TCP connection between a prod ECS task and the shared-services VPC endpoint ENI - TGW flow logs enabled for connectivity troubleshooting
PrivateLink checklist
- MCP server behind an internal NLB — Application Load Balancer not supported as PrivateLink target
- VPC Endpoint Service created with
--acceptance-requiredfor external consumers - Allowed principals list explicitly names consumer account ARNs — wildcard principal not used
- EventBridge rule alerts provider team on new endpoint connection requests
- NLB target group health checks passing — MCP server port responding with TCP or HTTP success
- Consumer-side endpoint SG allows outbound TCP on MCP server port to endpoint ENI private IPs
- Private DNS tested: consumer VPC DNS resolves service name to endpoint ENI private IP (not public IP)
- End-to-end test: consumer ECS task calls MCP server via PrivateLink and receives valid response
Network Firewall checklist
- Dedicated firewall subnet created in each AZ (separate CIDR from private and public subnets)
- NFW deployed with one subnet mapping per AZ — NFW endpoint IDs recorded per AZ
- All three route tables configured: private subnet → NFW endpoint; firewall subnet → IGW; IGW edge → NFW endpoint (per AZ)
- IGW edge route table attached to internet gateway via
aws ec2 associate-route-table --gateway-id - Domain list rule group uses both
HTTP_HOSTandTLS_SNItarget types for HTTPS destinations - Firewall policy default stateful action set to
aws:alert_strictduring observation period - Alert logs flowing to CloudWatch Logs — verified log group has recent entries
- 24–48 hour observation window completed; all legitimate destinations identified from alert logs and added to domain list
- Policy updated to
aws:drop_strictafter observation — confirmed no legitimate traffic blocked (zero unexpected connection failures for 1 hour post-switch) - Gateway endpoints for S3 and DynamoDB verified still in route table after NFW routing changes (prefix list routes take priority over
0.0.0.0/0)
Keep your MCP servers monitored while you secure the network
Locking down VPC routing, tightening endpoint policies, and deploying Network Firewall reduces attack surface — but it also introduces new failure modes. A misconfigured route table silently breaks Bedrock calls; a missing ECR endpoint causes deployment failures at 2am. AliveMCP pings every MCP endpoint every 60 seconds and alerts you before your users notice. See plans →