AWS Route 53 · 2026-10-02 · Route 53 arc
AWS Route 53 for MCP Servers: DNS Architecture, Health Checks, Failover, and Private Service Discovery
Route 53 is the DNS layer that every production MCP server deployment depends on — and DNS is more consequential for MCP servers than for most HTTP APIs because MCP clients embed the server hostname in tool registries, agent configurations, and workspace settings that are not re-read on every request. A DNS misconfiguration that changes the hostname, introduces intermittent failures, or misroutes traffic to a degraded endpoint breaks every agent that has saved the URL — silently, from the operator's perspective, until user-facing errors surface. This guide synthesizes five Route 53 topics — hosted zones and Alias records, health checks, failover routing, latency-based routing, and private hosted zones — into three structural patterns that govern the full DNS lifecycle for MCP deployments: record type selection and TTL strategy, high availability via health-check-driven routing, and advanced routing for multi-region and VPC-internal service meshes.
TL;DR
- Alias vs CNAME: Use Alias A records (not CNAME) for ALB, CloudFront, and any AWS-managed endpoint. Alias records are free (no per-query charge), work at the zone apex, resolve in a single hop, and support
EvaluateTargetHealth. App Runner custom domains require CNAME because App Runner has no hosted zone ID. - TTL ladder: Lower TTL to 60s at T−24h before planned deployments, keep it at 60s during cutover, raise back to 300s after stability is confirmed. MCP clients that cache the resolved IP for the process lifetime ignore DNS TTL — design zero-downtime at the load balancer layer, not in DNS.
- Failover: Active-passive uses PRIMARY record with
HealthCheckId+ SECONDARY without. Total failover window: ~90s detection + ~60s TTL propagation = 3–4 minutes at defaults. UseRequestInterval=10(fast health checks) to cut detection to 30s. If both checks fail, Route 53 returns SECONDARY anyway — it never returns NXDOMAIN. - Health check gap: Route 53 probes at the HTTP layer only. It cannot validate MCP
initialize,tools/list, auth middleware, or tool execution. Run AliveMCP alongside Route 53 for JSON-RPC protocol validation. - Latency routing: Effective for US/EU/APAC gaps (>100ms differential). Routing only applies to new DNS resolutions — in-flight MCP sessions hold the IP for their full lifetime. For per-packet routing during sessions, use Global Accelerator instead.
- Private hosted zones:
PrivateZone: True+ VPC association gives VPC-scoped DNS for internal MCP microservices. BothenableDnsSupportandenableDnsHostnamesmust be enabled on the VPC. Use Cloud Map for dynamic ECS task registration.
Pattern 1 — DNS fundamentals and record types
Creating a hosted zone and delegating NS records
Every Route 53 DNS configuration starts with a public hosted zone. The zone creation call returns four NS records that you delegate to at your domain registrar — after delegation, Route 53 handles all DNS resolution for that domain:
import boto3, uuid
route53 = boto3.client("route53")
response = route53.create_hosted_zone(
Name="mcp-api.example.com",
CallerReference=str(uuid.uuid4()), # unique per create call
HostedZoneConfig={
"Comment": "Public zone for MCP server endpoints",
"PrivateZone": False,
},
)
hosted_zone_id = response["HostedZone"]["Id"]
ns_records = response["DelegationSet"]["NameServers"]
# Set these 4 NS values at your registrar
# e.g. ['ns-100.awsdns-12.com', 'ns-200.awsdns-34.net', ...]
NS propagation after registrar delegation takes 24–48 hours for global resolution. Use dig +trace mcp-api.example.com to verify that the delegation chain resolves correctly from root servers down through the AWS-assigned name servers.
Alias records vs CNAME: the decision tree for MCP server endpoints
Every AWS-managed endpoint — ALB, CloudFront, App Runner, S3 website, API Gateway — is expressed as an AWS hostname (abc123.us-east-1.awsapprunner.com, my-alb-1234.us-east-1.elb.amazonaws.com). Route 53 Alias records are the correct DNS record type for these, with one exception:
ALB_HOSTED_ZONE_IDS = {
"us-east-1": "Z35SXDOTRQ7X7K",
"us-east-2": "Z3AADJGX6KTTL2",
"us-west-2": "Z1H1FL5HABSF5",
"eu-west-1": "Z32O12XQLNTSW2",
"ap-southeast-1": "Z1LMS91P8CMLE5",
}
# Alias A record → ALB (correct for zone apex and subdomains)
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"AliasTarget": {
"DNSName": "my-alb-1234567890.us-east-1.elb.amazonaws.com",
"HostedZoneId": ALB_HOSTED_ZONE_IDS["us-east-1"],
"EvaluateTargetHealth": True,
}
}
}]}
)
# App Runner custom domain → CNAME (App Runner has no hosted zone ID)
# After associate_custom_domain(), use the returned CNAME target
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "CNAME",
"TTL": 300,
"ResourceRecords": [{"Value": "abc123xyz.us-east-1.awsapprunner.com"}],
}
}]}
)
The four dimensions of the Alias vs CNAME decision for MCP deployments:
- Zone apex constraint: You cannot use a CNAME at the zone apex (
example.comitself) — only at sub-subdomains. If your MCP server lives at the root of a domain, use an Alias A record. App Runner is the main exception where CNAME is required for custom domains regardless of position. - Query cost: Route 53 does not charge for Alias record resolution when the Alias target is an AWS endpoint in the same account. CNAME queries bill at standard rates ($0.40 per million). For high-traffic MCP deployments queried repeatedly by agent runtimes, this adds up.
- Resolution hops: CNAME adds a second DNS lookup — the resolver must first resolve the CNAME target before getting an IP. Alias records are resolved in a single step by Route 53, returning the current ALB IPs directly to the resolver without an extra round-trip.
- Health evaluation: Setting
EvaluateTargetHealth: Trueon an Alias A record causes Route 53 to check whether the underlying AWS target is healthy (e.g., at least one ALB target group is healthy) before returning the record. This is free and requires no separate health check resource.
App Runner custom domain and ACM certificate validation
App Runner's associate_custom_domain() flow creates two sets of Route 53 records: one CNAME pointing at the App Runner endpoint, and one or more CNAME records for ACM certificate domain control validation:
apprunner = boto3.client("apprunner", region_name="us-east-1")
assoc = apprunner.associate_custom_domain(
ServiceArn="arn:aws:apprunner:us-east-1:123456789012:service/mcp-prod/...",
DomainName="mcp.example.com",
EnableWWWSubdomain=False,
)
# Step 1: CNAME for the App Runner endpoint itself
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "CNAME",
"TTL": 300,
"ResourceRecords": [{"Value": assoc["CustomDomain"]["DomainName"]}],
}
}]}
)
# Step 2: ACM validation CNAMEs — create all PENDING_VALIDATION records
for record in assoc["CustomDomain"]["CertificateValidationRecords"]:
if record["Status"] == "PENDING_VALIDATION":
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": record["Name"],
"Type": "CNAME",
"TTL": 300,
"ResourceRecords": [{"Value": record["Value"]}],
}
}]}
)
A critical operational detail: do not delete the ACM validation CNAME records after certificate issuance. ACM re-validates certificates on renewal (every 13 months) using the same CNAME values. Deleting them after the initial validation breaks automatic renewal silently — the certificate expires 13 months later without warning.
TTL ladder for MCP server deployments
TTL (Time To Live) controls how long resolvers cache a record before re-querying Route 53. For MCP server operators, TTL is a deployment velocity lever — low during planned changes, normal during stable operation:
# T-24h: lower TTL before any planned deployment
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "CNAME",
"TTL": 60, # ← drop to 60s pre-deployment
"ResourceRecords": [{"Value": "old-endpoint.example.com"}],
}
}]}
)
# At cutover: update the record — max propagation lag is now ~60s
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "CNAME",
"TTL": 60, # keep low until confirmed stable
"ResourceRecords": [{"Value": "new-endpoint.example.com"}],
}
}]}
)
# T+30min: raise TTL back after stable confirmation
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "CNAME",
"TTL": 300, # ← back to 5 min after stability confirmed
"ResourceRecords": [{"Value": "new-endpoint.example.com"}],
}
}]}
)
An important constraint for MCP deployments: many MCP client frameworks (Claude Desktop, IDE extensions, LLM agent runtimes) cache the resolved IP for the process lifetime and do not re-resolve DNS on every request. This means DNS TTL is not your only propagation concern — long-running agent processes hold a stale IP for hours regardless of TTL. Design zero-downtime deployments at the load balancer layer (rolling ECS task replacements, App Runner rolling service updates) rather than relying on DNS cutover to drain connections gracefully.
Pattern 2 — High availability: health checks and failover routing
Creating HTTPS health checks for MCP server endpoints
Route 53 health checks probe your MCP server endpoint from multiple AWS edge locations, mark it healthy or unhealthy, and feed that status directly into DNS routing decisions. Three parameters require careful tuning for MCP server workloads:
import uuid
health_check = route53.create_health_check(
CallerReference=str(uuid.uuid4()),
HealthCheckConfig={
"Type": "HTTPS",
# Must match the hostname in the TLS certificate
# Route 53 probers send this in the Host header and verify TLS against it
"FullyQualifiedDomainName": "mcp.example.com",
"Port": 443,
# Path must return HTTP 2xx — use a dedicated /health endpoint
"ResourcePath": "/health",
# 30 = standard (15 probers × 1 probe/30s = one probe every 2s globally)
# 10 = fast (costs ~3× more: $3/mo vs $0.75/mo per check)
"RequestInterval": 30,
# Consecutive failures before flipping to Unhealthy
# 3 × 30s = 90 seconds to detection at standard rate
"FailureThreshold": 3,
# Must be True for HTTPS on production — validates TLS cert chain
"EnableSNI": True,
"Regions": ["us-east-1", "eu-west-1", "ap-southeast-1"],
},
)
health_check_id = health_check["HealthCheck"]["Id"]
# Tag for identification
route53.change_tags_for_resource(
ResourceType="healthcheck",
ResourceId=health_check_id,
AddTags=[
{"Key": "Service", "Value": "mcp-server-prod"},
{"Key": "Environment", "Value": "production"},
],
)
For MCP servers behind an ALB, EvaluateTargetHealth: True on the Alias record handles load balancer health automatically — no explicit Route 53 health check is needed unless you also want SNS alerting independent of failover routing. For servers not behind an AWS-managed load balancer (bare EC2, on-premises, non-ALB), an explicit health check is the only way to get DNS-level failover.
CloudWatch alarms for health check failures
Route 53 publishes health check status as a CloudWatch metric — but only in us-east-1, regardless of where the probed endpoint is hosted. Always create Route 53-related CloudWatch alarms in us-east-1:
cloudwatch = boto3.client("cloudwatch", region_name="us-east-1")
sns = boto3.client("sns", region_name="us-east-1")
topic_arn = sns.create_topic(Name="mcp-server-route53-alerts")["TopicArn"]
sns.subscribe(TopicArn=topic_arn, Protocol="email", Endpoint="oncall@example.com")
cloudwatch.put_metric_alarm(
AlarmName="mcp-server-route53-healthcheck",
MetricName="HealthCheckStatus",
Namespace="AWS/Route53",
Dimensions=[{"Name": "HealthCheckId", "Value": health_check_id}],
Statistic="Minimum",
Period=60,
EvaluationPeriods=1,
Threshold=1.0,
ComparisonOperator="LessThanThreshold",
AlarmActions=[topic_arn],
OKActions=[topic_arn],
TreatMissingData="breaching",
)
What Route 53 health checks do and do not verify for MCP servers
Understanding the gap between Route 53 health checks and full MCP protocol validation is essential for building a complete monitoring stack:
# Route 53 health checks VERIFY:
# ✓ TCP port accepts connections
# ✓ HTTP response status is 2xx
# ✓ TLS certificate is valid and not expired (HTTPS checks with EnableSNI)
# ✓ Response body contains a search string (optional, first 5120 bytes)
# ✓ Response time within probe timeout (10s default)
# Route 53 health checks DO NOT verify:
# ✗ JSON-RPC initialize request succeeds
# ✗ tools/list returns a valid schema with expected tools
# ✗ Individual MCP tools execute correctly
# ✗ Auth (OAuth tokens, API keys) is accepted
# ✗ Schema drift — tools/list changed since last deployment
# ✗ P95/P99 latency at the MCP tool layer
# ✗ Transport-specific issues (SSE reconnects, streamable HTTP)
# This /health endpoint passes Route 53 checks even when:
# - tools/list returns an empty list (tool registry failed to load)
# - DB is down (all data-access tools fail silently)
# - Auth middleware is misconfigured (authenticated calls fail)
# - Server is in a half-initialized state after a crash-loop restart
@app.get("/health")
async def health():
return {"status": "ok"}
Route 53 health checks are the right tool for DNS failover — they detect infrastructure-level failures fast enough to update DNS routing. For MCP protocol-level validation — verifying that the server properly handles initialize, tools/list, and authenticated tool calls — use AliveMCP, which probes at the JSON-RPC layer and alerts on schema drift, tool-call failures, and authentication breakage that a simple HTTP health check misses entirely.
Active-passive failover: the DNS-layer circuit breaker
Active-passive failover keeps one endpoint live (primary) and one on standby (secondary). Route 53 automatically routes DNS to the secondary when the primary health check fails — no manual intervention required:
primary_hc = route53.create_health_check(
CallerReference=str(uuid.uuid4()),
HealthCheckConfig={
"Type": "HTTPS",
"FullyQualifiedDomainName": "mcp.example.com",
"Port": 443,
"ResourcePath": "/health",
"RequestInterval": 30,
"FailureThreshold": 3, # 90s to failover
"EnableSNI": True,
},
)
primary_hc_id = primary_hc["HealthCheck"]["Id"]
# PRIMARY record — withheld from DNS responses when health check is unhealthy
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": "mcp-primary-us-east-1",
"Failover": "PRIMARY",
"TTL": 60,
"ResourceRecords": [{"Value": "203.0.113.10"}],
"HealthCheckId": primary_hc_id,
}
}]}
)
# SECONDARY record — returned only when primary health check fails
# If secondary also has a health check and it fails, Route 53 still returns
# secondary — it will never return NXDOMAIN, even when all checks fail
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": "mcp-secondary-us-west-2",
"Failover": "SECONDARY",
"TTL": 60,
"ResourceRecords": [{"Value": "203.0.113.20"}],
}
}]}
)
Failover timing at standard settings: 3 consecutive failures at 30s interval = 90 seconds to detection, plus up to 60 seconds for cached records to expire. Total failover window from failure to all-new-clients-redirected: approximately 3–4 minutes. Using RequestInterval=10 (fast health checks) cuts detection time to 30 seconds at roughly 3× the cost per health check per month ($3.00 vs $0.75).
Active-active failover with weighted routing
Active-active distributes traffic across multiple MCP endpoints in normal operation, and removes a failed endpoint from rotation automatically when its health check fails:
endpoints = [
{"id": "us-east-1", "ip": "203.0.113.10", "weight": 50},
{"id": "us-west-2", "ip": "203.0.113.20", "weight": 50},
]
for endpoint in endpoints:
hc = route53.create_health_check(
CallerReference=str(uuid.uuid4()),
HealthCheckConfig={
"Type": "HTTPS",
"IPAddress": endpoint["ip"],
"FullyQualifiedDomainName": "mcp.example.com",
"Port": 443,
"ResourcePath": "/health",
"RequestInterval": 30,
"FailureThreshold": 3,
"EnableSNI": True,
},
)
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": f"mcp-active-{endpoint['id']}",
"Weight": endpoint["weight"],
"TTL": 60,
"ResourceRecords": [{"Value": endpoint["ip"]}],
"HealthCheckId": hc["HealthCheck"]["Id"],
}
}]}
)
With active-active weighted routing: during normal operation resolvers see both records and split traffic. When one endpoint fails, Route 53 omits it and the surviving endpoint absorbs 100% of traffic. When both endpoints fail simultaneously, Route 53 returns both records to avoid NXDOMAIN — clients will fail at TCP/TLS rather than DNS, which is the lesser evil. If complete-outage protection is more important than load distribution, prefer active-passive over active-active.
Manual failover for planned maintenance
Route 53 lets you force a health check into an unhealthy state without requiring an actual endpoint failure — the safest mechanism for planned maintenance on the primary MCP server:
import time
def failover_to_secondary(health_check_id: str):
route53.update_health_check(
HealthCheckId=health_check_id,
Disabled=True, # always evaluates to unhealthy
)
print("Primary health check disabled. DNS updates within TTL (~60s).")
def restore_primary(health_check_id: str):
route53.update_health_check(
HealthCheckId=health_check_id,
Disabled=False,
)
# Poll until 70% of probers report healthy
deadline = time.time() + 300
while time.time() < deadline:
response = route53.get_health_check_status(HealthCheckId=health_check_id)
statuses = [obs["StatusReport"]["Status"]
for obs in response["HealthCheckObservations"]]
healthy_count = sum(1 for s in statuses if s.startswith("Success"))
if healthy_count >= len(statuses) * 0.7:
print("Primary endpoint healthy — failback complete.")
return
time.sleep(30)
# Maintenance workflow:
# 1. failover_to_secondary() → traffic moves to secondary via DNS
# 2. Perform maintenance on primary
# 3. restore_primary() → monitors until healthy, then DNS routes back automatically
MCP-specific failover edge cases
Three behaviors matter specifically for MCP server operators:
- In-flight sessions during failover: Route 53 DNS failover only redirects new DNS resolutions. MCP clients that have already resolved the primary IP and established a session continue on that IP until the TCP connection drops or the session ends. DNS failover does not terminate existing sessions — design MCP server graceful shutdown to close SSE streams cleanly on SIGTERM so clients reconnect quickly and re-resolve DNS to the new endpoint.
- Both-fail behavior: When both primary and secondary health checks fail, Route 53 returns the SECONDARY record rather than NXDOMAIN. This is intentional — a complete DNS blackout is worse than returning a known-unhealthy endpoint. For write-critical MCP servers, prefer a secondary that returns clear JSON-RPC error responses over a secondary that is simply unavailable.
- Failback flapping: Route 53 restores the primary record after the primary health check passes
FailureThresholdconsecutive intervals. Brief primary flaps (30s failures between healthy periods) can cause repeated failover-failback cycles. Use AliveMCP's minimum-stable-duration alert suppression to avoid alert fatigue during flapping events.
Pattern 3 — Advanced routing: latency-based multi-region and private service discovery
Latency-based routing for multi-region MCP deployments
Route 53 latency-based routing directs each DNS query to the AWS region with the lowest measured round-trip time from the resolver's location. For MCP servers, latency routing matters because sessions are long-lived — a session running 30 tool calls over 10 minutes accumulates per-call latency into measurable user experience differences:
REGIONAL_ENDPOINTS = {
"us-east-1": {
"alb_dns": "mcp-us-east-1.us-east-1.elb.amazonaws.com",
"alb_zone_id": "Z35SXDOTRQ7X7K",
"ip": "203.0.113.10",
},
"eu-west-1": {
"alb_dns": "mcp-eu-west-1.eu-west-1.elb.amazonaws.com",
"alb_zone_id": "Z32O12XQLNTSW2",
"ip": "203.0.113.11",
},
"ap-southeast-1": {
"alb_dns": "mcp-apac.ap-southeast-1.elb.amazonaws.com",
"alb_zone_id": "Z1LMS91P8CMLE5",
"ip": "203.0.113.12",
},
}
for region, endpoint in REGIONAL_ENDPOINTS.items():
hc = route53.create_health_check(
CallerReference=str(uuid.uuid4()),
HealthCheckConfig={
"Type": "HTTPS",
"FullyQualifiedDomainName": "mcp.example.com",
"IPAddress": endpoint["ip"],
"Port": 443,
"ResourcePath": "/health",
"RequestInterval": 30,
"FailureThreshold": 3,
"EnableSNI": True,
"Regions": [region],
},
)
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": f"mcp-latency-{region}",
"Region": region, # the AWS region this record routes to
"AliasTarget": {
"DNSName": endpoint["alb_dns"],
"HostedZoneId": endpoint["alb_zone_id"],
"EvaluateTargetHealth": True,
},
"HealthCheckId": hc["HealthCheck"]["Id"],
}
}]}
)
When a regional health check fails, Route 53 excludes that region and routes queries to the next-lowest-latency available region automatically. The layered architecture (Route 53 latency record → regional ALB → ECS multi-AZ target groups) handles three failure domains independently: cross-region latency routing at the DNS layer, regional endpoint failover at the ALB layer, and AZ-level instance failures at the target group layer.
How Route 53 latency measurements work: the session stickiness constraint
Two non-obvious properties of Route 53 latency routing determine whether it's worth deploying for a given MCP workload:
# Route 53 latency routing — what it actually measures:
# Background AWS measurements from Route 53 edge probers to each region.
# Updated on the order of minutes — does NOT track transient network events.
# EDNS0 Client Subnet (ECS) used when available — routes based on actual
# client subnet rather than resolver datacenter location.
#
# What it works well for:
# - US users → us-east-1, EU users → eu-west-1 (>100ms gap, stable)
# - Asia-Pacific users → ap-southeast-1 (>150ms gap from US/EU)
#
# What it does NOT work for:
# - us-east-1 vs us-east-2 routing (gap too small, measurement noise too high)
# - Within-region AZ selection (use ALB target groups)
#
# Session stickiness trap:
# T=0: Agent resolves mcp.example.com → 203.0.113.10 (us-east-1)
# T=0: Agent establishes MCP session, ALL tool calls go to us-east-1
# T=N: Route 53 latency data shifts, new queries would go to eu-west-1
# T=N: EXISTING agent session is unaffected — still on us-east-1
#
# For session-level per-packet latency optimization, use Global Accelerator:
# Route 53 latency: free DNS routing, first-connection only, ~100ms detection lag
# Global Accelerator: ~$0.025/hr + $0.01/GB, per-packet routing via AWS backbone,
# instant failover, full-session benefit
# Verify actual latency differential before deploying multi-region:
# curl -o /dev/null -s -w "%{time_total}\n" https://mcp-us.example.com/health
# curl -o /dev/null -s -w "%{time_total}\n" https://mcp-eu.example.com/health
# If gap is <50ms from your primary user base, multi-region adds overhead for minimal gain
Geolocation routing for data sovereignty
When data residency requirements mandate that EU clients always hit eu-west-1 regardless of latency, use geolocation routing instead. A critical behavior difference: geolocation routing does not fall back to the default record on health check failure — a EU-matched record that fails its health check still routes EU clients to the EU endpoint, not to a US one:
# EU record — all European countries → eu-west-1 (GDPR data residency)
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": "mcp-geo-eu",
"GeoLocation": {"ContinentCode": "EU"},
"TTL": 60,
"ResourceRecords": [{"Value": "203.0.113.11"}],
"HealthCheckId": eu_hc_id,
}
}]}
)
# Default record — required; catches all countries not explicitly listed
# Without a default, clients from unlisted countries receive NXDOMAIN
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": "mcp-geo-default",
"GeoLocation": {"CountryCode": "*"}, # wildcard
"TTL": 60,
"ResourceRecords": [{"Value": "203.0.113.10"}],
"HealthCheckId": us_hc_id,
}
}]}
)
The no-fallback behavior is intentional for compliance — an EU client should never be routed to a US endpoint even during an outage, because the data sovereignty violation is worse than temporary unavailability. Design the EU secondary (within-region ALB failover) at the load balancer layer rather than relying on Route 53 cross-border fallback.
Private hosted zones for VPC-internal MCP services
Private hosted zones provide VPC-scoped DNS that resolves only from within associated VPCs — enabling internal MCP service discovery without exposing service hostnames to the public internet:
ec2 = boto3.client("ec2", region_name="us-east-1")
vpc_id = "vpc-0abc1234567890def"
# Create private hosted zone associated with the MCP server VPC
private_zone = route53.create_hosted_zone(
Name="mcp.internal", # .internal reserved — no TLD conflict
CallerReference=str(uuid.uuid4()),
HostedZoneConfig={
"Comment": "Internal DNS for MCP microservices",
"PrivateZone": True, # critical: makes it VPC-scoped
},
VPC={"VPCRegion": "us-east-1", "VPCId": vpc_id},
)
private_zone_id = private_zone["HostedZone"]["Id"]
# Register internal MCP service hostnames
services = [
("tool-executor.mcp.internal", "10.0.1.50"),
("auth-server.mcp.internal", "10.0.1.51"),
("registry-db.mcp.internal", "10.0.2.100"),
("mcp-gateway.mcp.internal", "10.0.1.10"),
]
for hostname, private_ip in services:
route53.change_resource_record_sets(
HostedZoneId=private_zone_id,
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": hostname,
"Type": "A",
"TTL": 60,
"ResourceRecords": [{"Value": private_ip}],
}
}]}
)
# VPC prerequisites — both must be enabled:
ec2.modify_vpc_attribute(
VpcId=vpc_id,
EnableDnsSupport={"Value": True}, # activates VPC resolver at CIDR+2
)
ec2.modify_vpc_attribute(
VpcId=vpc_id,
EnableDnsHostnames={"Value": True}, # assigns DNS hostnames to instances
)
# VPC resolver address: VPC_CIDR+2 (e.g., 10.0.0.2 for 10.0.0.0/16)
Split-horizon DNS for MCP dev/prod config parity
Split-horizon DNS uses the same hostname in both development and production environments but resolves it to different endpoints depending on which VPC the query comes from. This eliminates environment-specific configuration in MCP server code:
# Both dev and prod use "db.mcp.internal" — the VPC resolver selects the right target
# DATABASE_URL = "postgresql://db.mcp.internal:5432/mcp_db" # same in all envs
# Production private zone — prod VPC resolves to RDS proxy for connection pooling
prod_zone = route53.create_hosted_zone(
Name="mcp.internal", CallerReference=str(uuid.uuid4()),
HostedZoneConfig={"PrivateZone": True},
VPC={"VPCRegion": "us-east-1", "VPCId": "vpc-prod"},
)
route53.change_resource_record_sets(
HostedZoneId=prod_zone["HostedZone"]["Id"],
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "db.mcp.internal",
"Type": "CNAME",
"TTL": 60,
"ResourceRecords": [{"Value": "mcp-db-prod.proxy-abc123.us-east-1.rds.amazonaws.com"}],
}
}]}
)
# Dev private zone — same name, different VPC, resolves to dev RDS
dev_zone = route53.create_hosted_zone(
Name="mcp.internal", CallerReference=str(uuid.uuid4()),
HostedZoneConfig={"PrivateZone": True},
VPC={"VPCRegion": "us-east-1", "VPCId": "vpc-dev"},
)
route53.change_resource_record_sets(
HostedZoneId=dev_zone["HostedZone"]["Id"],
ChangeBatch={"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "db.mcp.internal",
"Type": "CNAME",
"TTL": 60,
"ResourceRecords": [{"Value": "mcp-db-dev.cdef5678.us-east-1.rds.amazonaws.com"}],
}
}]}
)
Cloud Map ECS Service Discovery for dynamic MCP microservices
AWS Cloud Map automatically registers ECS task private IPs as Route 53 records when tasks start and deregisters them when tasks stop — providing DNS for MCP microservices without manual IP management:
servicediscovery = boto3.client("servicediscovery", region_name="us-east-1")
ecs = boto3.client("ecs", region_name="us-east-1")
# Create Cloud Map private DNS namespace — auto-creates a Route 53 private zone
namespace = servicediscovery.create_private_dns_namespace(
Name="mcp.internal",
Vpc="vpc-0abc1234567890def",
Description="Service discovery for MCP microservices",
)
# Poll until creation completes, then get namespace_id...
# Create a Cloud Map service — defines the DNS record type and routing policy
sd_service = servicediscovery.create_service(
Name="tool-executor",
NamespaceId=namespace_id,
DnsConfig={
"DnsRecords": [{"Type": "A", "TTL": 10}], # low TTL for fast task-fail failover
"RoutingPolicy": "MULTIVALUE", # returns up to 8 healthy task IPs
},
HealthCheckCustomConfig={"FailureThreshold": 1},
)
# Create ECS Fargate service with Service Discovery enabled
ecs.create_service(
cluster="mcp-cluster",
serviceName="tool-executor",
taskDefinition="tool-executor:latest",
desiredCount=3,
launchType="FARGATE",
networkConfiguration={
"awsvpcConfiguration": {
"subnets": ["subnet-private-1a", "subnet-private-1b"],
"securityGroups": ["sg-tool-executor"],
"assignPublicIp": "DISABLED",
}
},
serviceRegistries=[{"registryArn": sd_service["Service"]["Arn"]}],
)
# Result: ECS tasks register as tool-executor.mcp.internal → [10.0.1.50, 10.0.1.51, ...]
# On task termination: deregistered within seconds
# MCP gateway resolves the name to get current healthy task IPs
Cloud Map's MULTIVALUE routing returns up to 8 healthy task IPs per DNS response. The MCP gateway connecting to tool-executor.mcp.internal should implement a retry loop that tries each returned IP sequentially if one fails — this is the client-side load balancing pattern for Cloud Map service discovery. Set TTL to 10s (not the default 30s) so that task replacements during rolling deployments propagate quickly to the gateway.
Failure modes reference table
| Failure | Symptom | Root cause | Fix |
|---|---|---|---|
| CNAME at zone apex | Record creation fails: "InvalidChangeBatch" | CNAME not permitted at zone apex in DNS spec | Use Alias A record instead of CNAME for root domain |
| App Runner TLS expiry 13 months after launch | HTTPS failures with cert-expired error | ACM validation CNAME deleted after initial issuance | Restore ACM validation CNAMEs returned by associate_custom_domain(); never delete them |
| Health check SSL failure | HealthCheckStatus alternates unhealthy/healthy | FullyQualifiedDomainName doesn't match TLS cert CN | Set FQDN to the exact hostname in the cert SAN; set EnableSNI=True |
| CloudWatch alarm creation fails | "ResourceNotFoundException: HealthCheck not found" | Route 53 metrics only exist in us-east-1; alarm created in wrong region | Always create Route 53 health check CloudWatch alarms in us-east-1 |
| Failover not triggering despite failures | DNS still returns PRIMARY during confirmed outage | HealthCheckId not attached to PRIMARY record; or TTL too high | Verify HealthCheckId field is populated on PRIMARY record; lower TTL to 60s |
| Active MCP sessions unaffected by failover | Users on established sessions keep hitting failed primary | DNS failover only redirects new resolutions; existing sessions hold cached IP | Graceful shutdown: close SSE streams on SIGTERM so clients reconnect and re-resolve |
| Both-fail returns stale secondary | DNS returns secondary during total outage; users get unexpected endpoint | Route 53 returns SECONDARY rather than NXDOMAIN when all checks fail | Configure secondary to return clear JSON-RPC errors instead of silent failures |
| Latency routing sends EU clients to US | Users in Germany connecting to us-east-1 | Latency routing uses resolver IP, not client IP; resolver in a US-based cloud | For compliance use geolocation routing; for performance confirm EDNS0 ECS is forwarded |
| Geolocation not falling back on health check failure | EU clients get NXDOMAIN during eu-west-1 outage | Geolocation routing does not cross borders on health check failure by design | Add secondary failover within eu-west-1 at ALB layer; accept compliance constraint |
| Private zone not resolving | dig mcp.internal returns NXDOMAIN from within VPC | enableDnsSupport or enableDnsHostnames not enabled on VPC | modify_vpc_attribute to enable both DNS settings; resolver is at VPC_CIDR+2 |
| Cloud Map task IP not deregistering | Dead ECS task IP remains in DNS responses after task termination | HealthCheckCustomConfig.FailureThreshold too high; Cloud Map deregistration delayed | Set FailureThreshold=1; implement client-side retry on connection failure to next IP |
| Route 53 health check probing wrong endpoint during ALB failover | Health check passes but ALB targets are all unhealthy | Separate Route 53 health check probes a healthy path not backed by ALB targets | Use EvaluateTargetHealth=True on Alias records + explicit health check for belt-and-suspenders |
Production checklists
Public hosted zone and record setup
- Hosted zone created; 4 NS records delegated at registrar
- Alias A record used for ALB / CloudFront / API Gateway endpoints (not CNAME)
- CNAME used for App Runner custom domain (no Alias option — App Runner has no hosted zone ID)
- ACM validation CNAMEs created and permanently retained (never deleted)
- TTL set to 60s on all MCP endpoint records (not 300s default — enable fast cutover)
- EvaluateTargetHealth=True on all Alias records pointing at ALBs
Health checks and failover
- HTTPS health check with EnableSNI=True, FullyQualifiedDomainName matching TLS cert CN
- ResourcePath pointing at a dedicated
/healthendpoint (not the MCP JSON-RPC path) - FailureThreshold=3, RequestInterval=30 (standard); use RequestInterval=10 only if 90s detection is too slow
- HealthCheckId attached to PRIMARY failover record (not just created — must be linked)
- SECONDARY record exists; if SECONDARY also has health check, document the both-fail behavior
- CloudWatch alarm on HealthCheckStatus metric created in us-east-1 regardless of endpoint region
- SNS topic and email/PagerDuty subscription wired to alarm AlarmActions and OKActions
- AliveMCP configured alongside Route 53 for JSON-RPC protocol-level monitoring
Multi-region and advanced routing
- Latency records created with one SetIdentifier per region; health check per region
- Regional ALB in each region; ECS multi-AZ target groups for within-region HA
- Geolocation default record present (CountryCode: "*") — required to avoid NXDOMAIN for unlisted countries
- Data sovereignty routing validated: EU queries hit EU region regardless of health check state
Private hosted zones and service discovery
- PrivateZone=True set at creation; VPC association confirmed
- VPC has enableDnsSupport=True and enableDnsHostnames=True
- Private zone uses .internal suffix (avoids TLD conflicts, not .local which conflicts with mDNS)
- Cloud Map service created with MULTIVALUE routing for ECS task IPs; TTL=10s
- MCP gateway client implements retry loop across all IPs returned in Cloud Map DNS response
- Split-horizon validated: same hostname resolves differently in dev vs prod VPC
Route 53 routes traffic — AliveMCP validates what receives it
Route 53 health checks tell you when your MCP server's HTTP port is closed. AliveMCP tells you when the port is open but the MCP protocol is broken — failed initialize handshakes, empty tools/list responses, auth middleware failures, schema drift after deployments. Route 53 handles DNS-layer failover; AliveMCP covers the protocol layer that Route 53 cannot reach. Run both for complete MCP server observability.