AWS Route 53 · 2026-10-02 · Route 53 arc

AWS Route 53 for MCP Servers: DNS Architecture, Health Checks, Failover, and Private Service Discovery

Route 53 is the DNS layer that every production MCP server deployment depends on — and DNS is more consequential for MCP servers than for most HTTP APIs because MCP clients embed the server hostname in tool registries, agent configurations, and workspace settings that are not re-read on every request. A DNS misconfiguration that changes the hostname, introduces intermittent failures, or misroutes traffic to a degraded endpoint breaks every agent that has saved the URL — silently, from the operator's perspective, until user-facing errors surface. This guide synthesizes five Route 53 topics — hosted zones and Alias records, health checks, failover routing, latency-based routing, and private hosted zones — into three structural patterns that govern the full DNS lifecycle for MCP deployments: record type selection and TTL strategy, high availability via health-check-driven routing, and advanced routing for multi-region and VPC-internal service meshes.

TL;DR

Pattern 1 — DNS fundamentals and record types

Creating a hosted zone and delegating NS records

Every Route 53 DNS configuration starts with a public hosted zone. The zone creation call returns four NS records that you delegate to at your domain registrar — after delegation, Route 53 handles all DNS resolution for that domain:

import boto3, uuid

route53 = boto3.client("route53")

response = route53.create_hosted_zone(
    Name="mcp-api.example.com",
    CallerReference=str(uuid.uuid4()),  # unique per create call
    HostedZoneConfig={
        "Comment": "Public zone for MCP server endpoints",
        "PrivateZone": False,
    },
)

hosted_zone_id = response["HostedZone"]["Id"]
ns_records = response["DelegationSet"]["NameServers"]
# Set these 4 NS values at your registrar
# e.g. ['ns-100.awsdns-12.com', 'ns-200.awsdns-34.net', ...]

NS propagation after registrar delegation takes 24–48 hours for global resolution. Use dig +trace mcp-api.example.com to verify that the delegation chain resolves correctly from root servers down through the AWS-assigned name servers.

Alias records vs CNAME: the decision tree for MCP server endpoints

Every AWS-managed endpoint — ALB, CloudFront, App Runner, S3 website, API Gateway — is expressed as an AWS hostname (abc123.us-east-1.awsapprunner.com, my-alb-1234.us-east-1.elb.amazonaws.com). Route 53 Alias records are the correct DNS record type for these, with one exception:

ALB_HOSTED_ZONE_IDS = {
    "us-east-1":      "Z35SXDOTRQ7X7K",
    "us-east-2":      "Z3AADJGX6KTTL2",
    "us-west-2":      "Z1H1FL5HABSF5",
    "eu-west-1":      "Z32O12XQLNTSW2",
    "ap-southeast-1": "Z1LMS91P8CMLE5",
}

# Alias A record → ALB (correct for zone apex and subdomains)
route53.change_resource_record_sets(
    HostedZoneId=HOSTED_ZONE_ID,
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "mcp.example.com",
            "Type": "A",
            "AliasTarget": {
                "DNSName": "my-alb-1234567890.us-east-1.elb.amazonaws.com",
                "HostedZoneId": ALB_HOSTED_ZONE_IDS["us-east-1"],
                "EvaluateTargetHealth": True,
            }
        }
    }]}
)

# App Runner custom domain → CNAME (App Runner has no hosted zone ID)
# After associate_custom_domain(), use the returned CNAME target
route53.change_resource_record_sets(
    HostedZoneId=HOSTED_ZONE_ID,
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "mcp.example.com",
            "Type": "CNAME",
            "TTL": 300,
            "ResourceRecords": [{"Value": "abc123xyz.us-east-1.awsapprunner.com"}],
        }
    }]}
)

The four dimensions of the Alias vs CNAME decision for MCP deployments:

App Runner custom domain and ACM certificate validation

App Runner's associate_custom_domain() flow creates two sets of Route 53 records: one CNAME pointing at the App Runner endpoint, and one or more CNAME records for ACM certificate domain control validation:

apprunner = boto3.client("apprunner", region_name="us-east-1")

assoc = apprunner.associate_custom_domain(
    ServiceArn="arn:aws:apprunner:us-east-1:123456789012:service/mcp-prod/...",
    DomainName="mcp.example.com",
    EnableWWWSubdomain=False,
)

# Step 1: CNAME for the App Runner endpoint itself
route53.change_resource_record_sets(
    HostedZoneId=HOSTED_ZONE_ID,
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "mcp.example.com",
            "Type": "CNAME",
            "TTL": 300,
            "ResourceRecords": [{"Value": assoc["CustomDomain"]["DomainName"]}],
        }
    }]}
)

# Step 2: ACM validation CNAMEs — create all PENDING_VALIDATION records
for record in assoc["CustomDomain"]["CertificateValidationRecords"]:
    if record["Status"] == "PENDING_VALIDATION":
        route53.change_resource_record_sets(
            HostedZoneId=HOSTED_ZONE_ID,
            ChangeBatch={"Changes": [{
                "Action": "UPSERT",
                "ResourceRecordSet": {
                    "Name": record["Name"],
                    "Type": "CNAME",
                    "TTL": 300,
                    "ResourceRecords": [{"Value": record["Value"]}],
                }
            }]}
        )

A critical operational detail: do not delete the ACM validation CNAME records after certificate issuance. ACM re-validates certificates on renewal (every 13 months) using the same CNAME values. Deleting them after the initial validation breaks automatic renewal silently — the certificate expires 13 months later without warning.

TTL ladder for MCP server deployments

TTL (Time To Live) controls how long resolvers cache a record before re-querying Route 53. For MCP server operators, TTL is a deployment velocity lever — low during planned changes, normal during stable operation:

# T-24h: lower TTL before any planned deployment
route53.change_resource_record_sets(
    HostedZoneId=HOSTED_ZONE_ID,
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "mcp.example.com",
            "Type": "CNAME",
            "TTL": 60,   # ← drop to 60s pre-deployment
            "ResourceRecords": [{"Value": "old-endpoint.example.com"}],
        }
    }]}
)

# At cutover: update the record — max propagation lag is now ~60s
route53.change_resource_record_sets(
    HostedZoneId=HOSTED_ZONE_ID,
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "mcp.example.com",
            "Type": "CNAME",
            "TTL": 60,   # keep low until confirmed stable
            "ResourceRecords": [{"Value": "new-endpoint.example.com"}],
        }
    }]}
)

# T+30min: raise TTL back after stable confirmation
route53.change_resource_record_sets(
    HostedZoneId=HOSTED_ZONE_ID,
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "mcp.example.com",
            "Type": "CNAME",
            "TTL": 300,  # ← back to 5 min after stability confirmed
            "ResourceRecords": [{"Value": "new-endpoint.example.com"}],
        }
    }]}
)

An important constraint for MCP deployments: many MCP client frameworks (Claude Desktop, IDE extensions, LLM agent runtimes) cache the resolved IP for the process lifetime and do not re-resolve DNS on every request. This means DNS TTL is not your only propagation concern — long-running agent processes hold a stale IP for hours regardless of TTL. Design zero-downtime deployments at the load balancer layer (rolling ECS task replacements, App Runner rolling service updates) rather than relying on DNS cutover to drain connections gracefully.

Pattern 2 — High availability: health checks and failover routing

Creating HTTPS health checks for MCP server endpoints

Route 53 health checks probe your MCP server endpoint from multiple AWS edge locations, mark it healthy or unhealthy, and feed that status directly into DNS routing decisions. Three parameters require careful tuning for MCP server workloads:

import uuid

health_check = route53.create_health_check(
    CallerReference=str(uuid.uuid4()),
    HealthCheckConfig={
        "Type": "HTTPS",
        # Must match the hostname in the TLS certificate
        # Route 53 probers send this in the Host header and verify TLS against it
        "FullyQualifiedDomainName": "mcp.example.com",
        "Port": 443,
        # Path must return HTTP 2xx — use a dedicated /health endpoint
        "ResourcePath": "/health",
        # 30 = standard (15 probers × 1 probe/30s = one probe every 2s globally)
        # 10 = fast (costs ~3× more: $3/mo vs $0.75/mo per check)
        "RequestInterval": 30,
        # Consecutive failures before flipping to Unhealthy
        # 3 × 30s = 90 seconds to detection at standard rate
        "FailureThreshold": 3,
        # Must be True for HTTPS on production — validates TLS cert chain
        "EnableSNI": True,
        "Regions": ["us-east-1", "eu-west-1", "ap-southeast-1"],
    },
)
health_check_id = health_check["HealthCheck"]["Id"]

# Tag for identification
route53.change_tags_for_resource(
    ResourceType="healthcheck",
    ResourceId=health_check_id,
    AddTags=[
        {"Key": "Service", "Value": "mcp-server-prod"},
        {"Key": "Environment", "Value": "production"},
    ],
)

For MCP servers behind an ALB, EvaluateTargetHealth: True on the Alias record handles load balancer health automatically — no explicit Route 53 health check is needed unless you also want SNS alerting independent of failover routing. For servers not behind an AWS-managed load balancer (bare EC2, on-premises, non-ALB), an explicit health check is the only way to get DNS-level failover.

CloudWatch alarms for health check failures

Route 53 publishes health check status as a CloudWatch metric — but only in us-east-1, regardless of where the probed endpoint is hosted. Always create Route 53-related CloudWatch alarms in us-east-1:

cloudwatch = boto3.client("cloudwatch", region_name="us-east-1")
sns = boto3.client("sns", region_name="us-east-1")

topic_arn = sns.create_topic(Name="mcp-server-route53-alerts")["TopicArn"]
sns.subscribe(TopicArn=topic_arn, Protocol="email", Endpoint="oncall@example.com")

cloudwatch.put_metric_alarm(
    AlarmName="mcp-server-route53-healthcheck",
    MetricName="HealthCheckStatus",
    Namespace="AWS/Route53",
    Dimensions=[{"Name": "HealthCheckId", "Value": health_check_id}],
    Statistic="Minimum",
    Period=60,
    EvaluationPeriods=1,
    Threshold=1.0,
    ComparisonOperator="LessThanThreshold",
    AlarmActions=[topic_arn],
    OKActions=[topic_arn],
    TreatMissingData="breaching",
)

What Route 53 health checks do and do not verify for MCP servers

Understanding the gap between Route 53 health checks and full MCP protocol validation is essential for building a complete monitoring stack:

# Route 53 health checks VERIFY:
# ✓ TCP port accepts connections
# ✓ HTTP response status is 2xx
# ✓ TLS certificate is valid and not expired (HTTPS checks with EnableSNI)
# ✓ Response body contains a search string (optional, first 5120 bytes)
# ✓ Response time within probe timeout (10s default)

# Route 53 health checks DO NOT verify:
# ✗ JSON-RPC initialize request succeeds
# ✗ tools/list returns a valid schema with expected tools
# ✗ Individual MCP tools execute correctly
# ✗ Auth (OAuth tokens, API keys) is accepted
# ✗ Schema drift — tools/list changed since last deployment
# ✗ P95/P99 latency at the MCP tool layer
# ✗ Transport-specific issues (SSE reconnects, streamable HTTP)

# This /health endpoint passes Route 53 checks even when:
# - tools/list returns an empty list (tool registry failed to load)
# - DB is down (all data-access tools fail silently)
# - Auth middleware is misconfigured (authenticated calls fail)
# - Server is in a half-initialized state after a crash-loop restart
@app.get("/health")
async def health():
    return {"status": "ok"}

Route 53 health checks are the right tool for DNS failover — they detect infrastructure-level failures fast enough to update DNS routing. For MCP protocol-level validation — verifying that the server properly handles initialize, tools/list, and authenticated tool calls — use AliveMCP, which probes at the JSON-RPC layer and alerts on schema drift, tool-call failures, and authentication breakage that a simple HTTP health check misses entirely.

Active-passive failover: the DNS-layer circuit breaker

Active-passive failover keeps one endpoint live (primary) and one on standby (secondary). Route 53 automatically routes DNS to the secondary when the primary health check fails — no manual intervention required:

primary_hc = route53.create_health_check(
    CallerReference=str(uuid.uuid4()),
    HealthCheckConfig={
        "Type": "HTTPS",
        "FullyQualifiedDomainName": "mcp.example.com",
        "Port": 443,
        "ResourcePath": "/health",
        "RequestInterval": 30,
        "FailureThreshold": 3,  # 90s to failover
        "EnableSNI": True,
    },
)
primary_hc_id = primary_hc["HealthCheck"]["Id"]

# PRIMARY record — withheld from DNS responses when health check is unhealthy
route53.change_resource_record_sets(
    HostedZoneId=HOSTED_ZONE_ID,
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "mcp.example.com",
            "Type": "A",
            "SetIdentifier": "mcp-primary-us-east-1",
            "Failover": "PRIMARY",
            "TTL": 60,
            "ResourceRecords": [{"Value": "203.0.113.10"}],
            "HealthCheckId": primary_hc_id,
        }
    }]}
)

# SECONDARY record — returned only when primary health check fails
# If secondary also has a health check and it fails, Route 53 still returns
# secondary — it will never return NXDOMAIN, even when all checks fail
route53.change_resource_record_sets(
    HostedZoneId=HOSTED_ZONE_ID,
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "mcp.example.com",
            "Type": "A",
            "SetIdentifier": "mcp-secondary-us-west-2",
            "Failover": "SECONDARY",
            "TTL": 60,
            "ResourceRecords": [{"Value": "203.0.113.20"}],
        }
    }]}
)

Failover timing at standard settings: 3 consecutive failures at 30s interval = 90 seconds to detection, plus up to 60 seconds for cached records to expire. Total failover window from failure to all-new-clients-redirected: approximately 3–4 minutes. Using RequestInterval=10 (fast health checks) cuts detection time to 30 seconds at roughly 3× the cost per health check per month ($3.00 vs $0.75).

Active-active failover with weighted routing

Active-active distributes traffic across multiple MCP endpoints in normal operation, and removes a failed endpoint from rotation automatically when its health check fails:

endpoints = [
    {"id": "us-east-1", "ip": "203.0.113.10", "weight": 50},
    {"id": "us-west-2", "ip": "203.0.113.20", "weight": 50},
]

for endpoint in endpoints:
    hc = route53.create_health_check(
        CallerReference=str(uuid.uuid4()),
        HealthCheckConfig={
            "Type": "HTTPS",
            "IPAddress": endpoint["ip"],
            "FullyQualifiedDomainName": "mcp.example.com",
            "Port": 443,
            "ResourcePath": "/health",
            "RequestInterval": 30,
            "FailureThreshold": 3,
            "EnableSNI": True,
        },
    )
    route53.change_resource_record_sets(
        HostedZoneId=HOSTED_ZONE_ID,
        ChangeBatch={"Changes": [{
            "Action": "UPSERT",
            "ResourceRecordSet": {
                "Name": "mcp.example.com",
                "Type": "A",
                "SetIdentifier": f"mcp-active-{endpoint['id']}",
                "Weight": endpoint["weight"],
                "TTL": 60,
                "ResourceRecords": [{"Value": endpoint["ip"]}],
                "HealthCheckId": hc["HealthCheck"]["Id"],
            }
        }]}
    )

With active-active weighted routing: during normal operation resolvers see both records and split traffic. When one endpoint fails, Route 53 omits it and the surviving endpoint absorbs 100% of traffic. When both endpoints fail simultaneously, Route 53 returns both records to avoid NXDOMAIN — clients will fail at TCP/TLS rather than DNS, which is the lesser evil. If complete-outage protection is more important than load distribution, prefer active-passive over active-active.

Manual failover for planned maintenance

Route 53 lets you force a health check into an unhealthy state without requiring an actual endpoint failure — the safest mechanism for planned maintenance on the primary MCP server:

import time

def failover_to_secondary(health_check_id: str):
    route53.update_health_check(
        HealthCheckId=health_check_id,
        Disabled=True,  # always evaluates to unhealthy
    )
    print("Primary health check disabled. DNS updates within TTL (~60s).")


def restore_primary(health_check_id: str):
    route53.update_health_check(
        HealthCheckId=health_check_id,
        Disabled=False,
    )
    # Poll until 70% of probers report healthy
    deadline = time.time() + 300
    while time.time() < deadline:
        response = route53.get_health_check_status(HealthCheckId=health_check_id)
        statuses = [obs["StatusReport"]["Status"]
                    for obs in response["HealthCheckObservations"]]
        healthy_count = sum(1 for s in statuses if s.startswith("Success"))
        if healthy_count >= len(statuses) * 0.7:
            print("Primary endpoint healthy — failback complete.")
            return
        time.sleep(30)


# Maintenance workflow:
# 1. failover_to_secondary() → traffic moves to secondary via DNS
# 2. Perform maintenance on primary
# 3. restore_primary() → monitors until healthy, then DNS routes back automatically

MCP-specific failover edge cases

Three behaviors matter specifically for MCP server operators:

Pattern 3 — Advanced routing: latency-based multi-region and private service discovery

Latency-based routing for multi-region MCP deployments

Route 53 latency-based routing directs each DNS query to the AWS region with the lowest measured round-trip time from the resolver's location. For MCP servers, latency routing matters because sessions are long-lived — a session running 30 tool calls over 10 minutes accumulates per-call latency into measurable user experience differences:

REGIONAL_ENDPOINTS = {
    "us-east-1": {
        "alb_dns": "mcp-us-east-1.us-east-1.elb.amazonaws.com",
        "alb_zone_id": "Z35SXDOTRQ7X7K",
        "ip": "203.0.113.10",
    },
    "eu-west-1": {
        "alb_dns": "mcp-eu-west-1.eu-west-1.elb.amazonaws.com",
        "alb_zone_id": "Z32O12XQLNTSW2",
        "ip": "203.0.113.11",
    },
    "ap-southeast-1": {
        "alb_dns": "mcp-apac.ap-southeast-1.elb.amazonaws.com",
        "alb_zone_id": "Z1LMS91P8CMLE5",
        "ip": "203.0.113.12",
    },
}

for region, endpoint in REGIONAL_ENDPOINTS.items():
    hc = route53.create_health_check(
        CallerReference=str(uuid.uuid4()),
        HealthCheckConfig={
            "Type": "HTTPS",
            "FullyQualifiedDomainName": "mcp.example.com",
            "IPAddress": endpoint["ip"],
            "Port": 443,
            "ResourcePath": "/health",
            "RequestInterval": 30,
            "FailureThreshold": 3,
            "EnableSNI": True,
            "Regions": [region],
        },
    )

    route53.change_resource_record_sets(
        HostedZoneId=HOSTED_ZONE_ID,
        ChangeBatch={"Changes": [{
            "Action": "UPSERT",
            "ResourceRecordSet": {
                "Name": "mcp.example.com",
                "Type": "A",
                "SetIdentifier": f"mcp-latency-{region}",
                "Region": region,  # the AWS region this record routes to
                "AliasTarget": {
                    "DNSName": endpoint["alb_dns"],
                    "HostedZoneId": endpoint["alb_zone_id"],
                    "EvaluateTargetHealth": True,
                },
                "HealthCheckId": hc["HealthCheck"]["Id"],
            }
        }]}
    )

When a regional health check fails, Route 53 excludes that region and routes queries to the next-lowest-latency available region automatically. The layered architecture (Route 53 latency record → regional ALB → ECS multi-AZ target groups) handles three failure domains independently: cross-region latency routing at the DNS layer, regional endpoint failover at the ALB layer, and AZ-level instance failures at the target group layer.

How Route 53 latency measurements work: the session stickiness constraint

Two non-obvious properties of Route 53 latency routing determine whether it's worth deploying for a given MCP workload:

# Route 53 latency routing — what it actually measures:
# Background AWS measurements from Route 53 edge probers to each region.
# Updated on the order of minutes — does NOT track transient network events.
# EDNS0 Client Subnet (ECS) used when available — routes based on actual
# client subnet rather than resolver datacenter location.
#
# What it works well for:
# - US users → us-east-1, EU users → eu-west-1 (>100ms gap, stable)
# - Asia-Pacific users → ap-southeast-1 (>150ms gap from US/EU)
#
# What it does NOT work for:
# - us-east-1 vs us-east-2 routing (gap too small, measurement noise too high)
# - Within-region AZ selection (use ALB target groups)
#
# Session stickiness trap:
# T=0:  Agent resolves mcp.example.com → 203.0.113.10 (us-east-1)
# T=0:  Agent establishes MCP session, ALL tool calls go to us-east-1
# T=N:  Route 53 latency data shifts, new queries would go to eu-west-1
# T=N:  EXISTING agent session is unaffected — still on us-east-1
#
# For session-level per-packet latency optimization, use Global Accelerator:
# Route 53 latency: free DNS routing, first-connection only, ~100ms detection lag
# Global Accelerator: ~$0.025/hr + $0.01/GB, per-packet routing via AWS backbone,
#                     instant failover, full-session benefit

# Verify actual latency differential before deploying multi-region:
# curl -o /dev/null -s -w "%{time_total}\n" https://mcp-us.example.com/health
# curl -o /dev/null -s -w "%{time_total}\n" https://mcp-eu.example.com/health
# If gap is <50ms from your primary user base, multi-region adds overhead for minimal gain

Geolocation routing for data sovereignty

When data residency requirements mandate that EU clients always hit eu-west-1 regardless of latency, use geolocation routing instead. A critical behavior difference: geolocation routing does not fall back to the default record on health check failure — a EU-matched record that fails its health check still routes EU clients to the EU endpoint, not to a US one:

# EU record — all European countries → eu-west-1 (GDPR data residency)
route53.change_resource_record_sets(
    HostedZoneId=HOSTED_ZONE_ID,
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "mcp.example.com",
            "Type": "A",
            "SetIdentifier": "mcp-geo-eu",
            "GeoLocation": {"ContinentCode": "EU"},
            "TTL": 60,
            "ResourceRecords": [{"Value": "203.0.113.11"}],
            "HealthCheckId": eu_hc_id,
        }
    }]}
)

# Default record — required; catches all countries not explicitly listed
# Without a default, clients from unlisted countries receive NXDOMAIN
route53.change_resource_record_sets(
    HostedZoneId=HOSTED_ZONE_ID,
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "mcp.example.com",
            "Type": "A",
            "SetIdentifier": "mcp-geo-default",
            "GeoLocation": {"CountryCode": "*"},  # wildcard
            "TTL": 60,
            "ResourceRecords": [{"Value": "203.0.113.10"}],
            "HealthCheckId": us_hc_id,
        }
    }]}
)

The no-fallback behavior is intentional for compliance — an EU client should never be routed to a US endpoint even during an outage, because the data sovereignty violation is worse than temporary unavailability. Design the EU secondary (within-region ALB failover) at the load balancer layer rather than relying on Route 53 cross-border fallback.

Private hosted zones for VPC-internal MCP services

Private hosted zones provide VPC-scoped DNS that resolves only from within associated VPCs — enabling internal MCP service discovery without exposing service hostnames to the public internet:

ec2 = boto3.client("ec2", region_name="us-east-1")
vpc_id = "vpc-0abc1234567890def"

# Create private hosted zone associated with the MCP server VPC
private_zone = route53.create_hosted_zone(
    Name="mcp.internal",      # .internal reserved — no TLD conflict
    CallerReference=str(uuid.uuid4()),
    HostedZoneConfig={
        "Comment": "Internal DNS for MCP microservices",
        "PrivateZone": True,   # critical: makes it VPC-scoped
    },
    VPC={"VPCRegion": "us-east-1", "VPCId": vpc_id},
)
private_zone_id = private_zone["HostedZone"]["Id"]

# Register internal MCP service hostnames
services = [
    ("tool-executor.mcp.internal", "10.0.1.50"),
    ("auth-server.mcp.internal", "10.0.1.51"),
    ("registry-db.mcp.internal", "10.0.2.100"),
    ("mcp-gateway.mcp.internal", "10.0.1.10"),
]
for hostname, private_ip in services:
    route53.change_resource_record_sets(
        HostedZoneId=private_zone_id,
        ChangeBatch={"Changes": [{
            "Action": "UPSERT",
            "ResourceRecordSet": {
                "Name": hostname,
                "Type": "A",
                "TTL": 60,
                "ResourceRecords": [{"Value": private_ip}],
            }
        }]}
    )

# VPC prerequisites — both must be enabled:
ec2.modify_vpc_attribute(
    VpcId=vpc_id,
    EnableDnsSupport={"Value": True},     # activates VPC resolver at CIDR+2
)
ec2.modify_vpc_attribute(
    VpcId=vpc_id,
    EnableDnsHostnames={"Value": True},   # assigns DNS hostnames to instances
)
# VPC resolver address: VPC_CIDR+2 (e.g., 10.0.0.2 for 10.0.0.0/16)

Split-horizon DNS for MCP dev/prod config parity

Split-horizon DNS uses the same hostname in both development and production environments but resolves it to different endpoints depending on which VPC the query comes from. This eliminates environment-specific configuration in MCP server code:

# Both dev and prod use "db.mcp.internal" — the VPC resolver selects the right target
# DATABASE_URL = "postgresql://db.mcp.internal:5432/mcp_db"  # same in all envs

# Production private zone — prod VPC resolves to RDS proxy for connection pooling
prod_zone = route53.create_hosted_zone(
    Name="mcp.internal", CallerReference=str(uuid.uuid4()),
    HostedZoneConfig={"PrivateZone": True},
    VPC={"VPCRegion": "us-east-1", "VPCId": "vpc-prod"},
)
route53.change_resource_record_sets(
    HostedZoneId=prod_zone["HostedZone"]["Id"],
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "db.mcp.internal",
            "Type": "CNAME",
            "TTL": 60,
            "ResourceRecords": [{"Value": "mcp-db-prod.proxy-abc123.us-east-1.rds.amazonaws.com"}],
        }
    }]}
)

# Dev private zone — same name, different VPC, resolves to dev RDS
dev_zone = route53.create_hosted_zone(
    Name="mcp.internal", CallerReference=str(uuid.uuid4()),
    HostedZoneConfig={"PrivateZone": True},
    VPC={"VPCRegion": "us-east-1", "VPCId": "vpc-dev"},
)
route53.change_resource_record_sets(
    HostedZoneId=dev_zone["HostedZone"]["Id"],
    ChangeBatch={"Changes": [{
        "Action": "UPSERT",
        "ResourceRecordSet": {
            "Name": "db.mcp.internal",
            "Type": "CNAME",
            "TTL": 60,
            "ResourceRecords": [{"Value": "mcp-db-dev.cdef5678.us-east-1.rds.amazonaws.com"}],
        }
    }]}
)

Cloud Map ECS Service Discovery for dynamic MCP microservices

AWS Cloud Map automatically registers ECS task private IPs as Route 53 records when tasks start and deregisters them when tasks stop — providing DNS for MCP microservices without manual IP management:

servicediscovery = boto3.client("servicediscovery", region_name="us-east-1")
ecs = boto3.client("ecs", region_name="us-east-1")

# Create Cloud Map private DNS namespace — auto-creates a Route 53 private zone
namespace = servicediscovery.create_private_dns_namespace(
    Name="mcp.internal",
    Vpc="vpc-0abc1234567890def",
    Description="Service discovery for MCP microservices",
)
# Poll until creation completes, then get namespace_id...

# Create a Cloud Map service — defines the DNS record type and routing policy
sd_service = servicediscovery.create_service(
    Name="tool-executor",
    NamespaceId=namespace_id,
    DnsConfig={
        "DnsRecords": [{"Type": "A", "TTL": 10}],  # low TTL for fast task-fail failover
        "RoutingPolicy": "MULTIVALUE",  # returns up to 8 healthy task IPs
    },
    HealthCheckCustomConfig={"FailureThreshold": 1},
)

# Create ECS Fargate service with Service Discovery enabled
ecs.create_service(
    cluster="mcp-cluster",
    serviceName="tool-executor",
    taskDefinition="tool-executor:latest",
    desiredCount=3,
    launchType="FARGATE",
    networkConfiguration={
        "awsvpcConfiguration": {
            "subnets": ["subnet-private-1a", "subnet-private-1b"],
            "securityGroups": ["sg-tool-executor"],
            "assignPublicIp": "DISABLED",
        }
    },
    serviceRegistries=[{"registryArn": sd_service["Service"]["Arn"]}],
)

# Result: ECS tasks register as tool-executor.mcp.internal → [10.0.1.50, 10.0.1.51, ...]
# On task termination: deregistered within seconds
# MCP gateway resolves the name to get current healthy task IPs

Cloud Map's MULTIVALUE routing returns up to 8 healthy task IPs per DNS response. The MCP gateway connecting to tool-executor.mcp.internal should implement a retry loop that tries each returned IP sequentially if one fails — this is the client-side load balancing pattern for Cloud Map service discovery. Set TTL to 10s (not the default 30s) so that task replacements during rolling deployments propagate quickly to the gateway.

Failure modes reference table

Failure Symptom Root cause Fix
CNAME at zone apex Record creation fails: "InvalidChangeBatch" CNAME not permitted at zone apex in DNS spec Use Alias A record instead of CNAME for root domain
App Runner TLS expiry 13 months after launch HTTPS failures with cert-expired error ACM validation CNAME deleted after initial issuance Restore ACM validation CNAMEs returned by associate_custom_domain(); never delete them
Health check SSL failure HealthCheckStatus alternates unhealthy/healthy FullyQualifiedDomainName doesn't match TLS cert CN Set FQDN to the exact hostname in the cert SAN; set EnableSNI=True
CloudWatch alarm creation fails "ResourceNotFoundException: HealthCheck not found" Route 53 metrics only exist in us-east-1; alarm created in wrong region Always create Route 53 health check CloudWatch alarms in us-east-1
Failover not triggering despite failures DNS still returns PRIMARY during confirmed outage HealthCheckId not attached to PRIMARY record; or TTL too high Verify HealthCheckId field is populated on PRIMARY record; lower TTL to 60s
Active MCP sessions unaffected by failover Users on established sessions keep hitting failed primary DNS failover only redirects new resolutions; existing sessions hold cached IP Graceful shutdown: close SSE streams on SIGTERM so clients reconnect and re-resolve
Both-fail returns stale secondary DNS returns secondary during total outage; users get unexpected endpoint Route 53 returns SECONDARY rather than NXDOMAIN when all checks fail Configure secondary to return clear JSON-RPC errors instead of silent failures
Latency routing sends EU clients to US Users in Germany connecting to us-east-1 Latency routing uses resolver IP, not client IP; resolver in a US-based cloud For compliance use geolocation routing; for performance confirm EDNS0 ECS is forwarded
Geolocation not falling back on health check failure EU clients get NXDOMAIN during eu-west-1 outage Geolocation routing does not cross borders on health check failure by design Add secondary failover within eu-west-1 at ALB layer; accept compliance constraint
Private zone not resolving dig mcp.internal returns NXDOMAIN from within VPC enableDnsSupport or enableDnsHostnames not enabled on VPC modify_vpc_attribute to enable both DNS settings; resolver is at VPC_CIDR+2
Cloud Map task IP not deregistering Dead ECS task IP remains in DNS responses after task termination HealthCheckCustomConfig.FailureThreshold too high; Cloud Map deregistration delayed Set FailureThreshold=1; implement client-side retry on connection failure to next IP
Route 53 health check probing wrong endpoint during ALB failover Health check passes but ALB targets are all unhealthy Separate Route 53 health check probes a healthy path not backed by ALB targets Use EvaluateTargetHealth=True on Alias records + explicit health check for belt-and-suspenders

Production checklists

Public hosted zone and record setup

Health checks and failover

Multi-region and advanced routing

Private hosted zones and service discovery

Route 53 routes traffic — AliveMCP validates what receives it

Route 53 health checks tell you when your MCP server's HTTP port is closed. AliveMCP tells you when the port is open but the MCP protocol is broken — failed initialize handshakes, empty tools/list responses, auth middleware failures, schema drift after deployments. Route 53 handles DNS-layer failover; AliveMCP covers the protocol layer that Route 53 cannot reach. Run both for complete MCP server observability.

Join the waitlist →