Guide · AWS Route 53 · MCP Server Health Checks · DNS Failover
Route 53 Health Checks for MCP Server Endpoints
AWS Route 53 health checks probe your MCP server endpoint from multiple AWS edge locations and mark it healthy or unhealthy — feeding that status directly into DNS routing decisions. A Route 53 health check linked to a failover record set will automatically stop routing DNS to a failed endpoint, making it the DNS-layer circuit breaker for multi-endpoint MCP deployments. But Route 53 health checks have a fundamental limitation for MCP servers: they probe at the HTTP layer only — they verify that /health returns 200 OK, but they cannot verify that the MCP JSON-RPC protocol is working correctly, that tools/list returns a valid schema, or that an individual tool is executing correctly. Three parameters require careful configuration for MCP server workloads: RequestInterval (30s standard or 10s fast — fast costs 3× more per health check per month), FailureThreshold (number of consecutive failures before the check flips unhealthy — default 3, meaning ~90s detection lag at standard rate), and FullyQualifiedDomainName (must match the MCP server's TLS certificate CN to avoid SSL validation failures).
TL;DR
Create an HTTP health check with route53.create_health_check() targeting your MCP server's /health path on port 443. Set FailureThreshold=3 and RequestInterval=30 for standard monitoring, or RequestInterval=10 for fast-response failover. Link the health check ID to a Route 53 DNS record's HealthCheckId field to enable automatic DNS failover. Create an SNS topic and CloudWatch alarm on the HealthCheckStatus metric for alerting. Use AliveMCP alongside Route 53 health checks for JSON-RPC protocol validation that Route 53 cannot perform.
Creating an HTTPS health check for an MCP server endpoint
The Route 53 create_health_check() API creates the health check resource. The CallerReference must be unique per call — use a UUID:
import boto3
import uuid
route53 = boto3.client("route53")
# HTTPS health check for an MCP server endpoint
response = route53.create_health_check(
CallerReference=str(uuid.uuid4()),
HealthCheckConfig={
"Type": "HTTPS",
# Must match the hostname in the TLS certificate
# Route 53 probers send this in the Host header and verify TLS against it
"FullyQualifiedDomainName": "mcp.example.com",
"Port": 443,
# Path to probe — must return HTTP 2xx for check to pass
# Use a dedicated health endpoint, NOT the MCP endpoint itself
"ResourcePath": "/health",
# How often to probe from each Route 53 edge location
# "30" = standard (15 probers × 1 probe per 30s = one probe every 2s globally)
# "10" = fast (costs ~3× more — $3/mo vs $0.75/mo per check per region)
"RequestInterval": 30,
# Consecutive failures before flipping to Unhealthy
# 3 failures at 30s interval = 90s to detection
"FailureThreshold": 3,
# If True, Route 53 also validates the TLS certificate chain
# Must be True for HTTPS checks on production endpoints
# Set False only if using self-signed certs (dev/test)
"EnableSNI": True,
# Optional: check that the response body contains a string
# Route 53 checks the first 5120 bytes of the response body
# "SearchString": "\"status\":\"ok\"",
# Regions to probe from (subset of available regions)
# Default: all regions (15 probers globally)
# Restricting regions reduces cost but limits geographic coverage
"Regions": [
"us-east-1",
"eu-west-1",
"ap-southeast-1",
],
},
)
health_check_id = response["HealthCheck"]["Id"]
print(f"Health check ID: {health_check_id}")
# Tag the health check for identification in the console
route53.change_tags_for_resource(
ResourceType="healthcheck",
ResourceId=health_check_id,
AddTags=[
{"Key": "Service", "Value": "mcp-server-prod"},
{"Key": "Environment", "Value": "production"},
],
)
Health check types for different MCP server configurations
Route 53 offers four check types. The right choice depends on how the MCP server is deployed:
# Type: HTTP — plain HTTP on a non-TLS port
# Use for: MCP servers behind a TLS-terminating ALB where
# the health check probe can reach the non-TLS port directly
{
"Type": "HTTP",
"FullyQualifiedDomainName": "mcp.example.com",
"Port": 80,
"ResourcePath": "/health",
"RequestInterval": 30,
"FailureThreshold": 3,
}
# Type: HTTPS — TLS endpoint (recommended for production)
{
"Type": "HTTPS",
"FullyQualifiedDomainName": "mcp.example.com",
"Port": 443,
"ResourcePath": "/health",
"EnableSNI": True,
"RequestInterval": 30,
"FailureThreshold": 3,
}
# Type: TCP — checks that the port accepts a TCP connection
# Use for: MCP servers where an HTTP health path is not available
# Limitation: does not verify the application is responding correctly,
# only that the port is listening
{
"Type": "TCP",
"FullyQualifiedDomainName": "mcp.example.com",
"Port": 443,
"RequestInterval": 30,
"FailureThreshold": 3,
}
# Type: CALCULATED — aggregates other health checks
# Use for: MCP server clusters where healthy = majority of instances healthy
# A Calculated check is healthy if the specified number of child checks are healthy
{
"Type": "CALCULATED",
"HealthThreshold": 2, # at least 2 of the child checks must be healthy
"ChildHealthChecks": [
"check-id-instance-1",
"check-id-instance-2",
"check-id-instance-3",
],
}
For MCP servers behind an ALB, the EvaluateTargetHealth flag on the Alias record handles load balancer health automatically. Create an explicit Route 53 health check when you need SNS alerting or when the MCP endpoint is not behind an AWS-managed load balancer.
Linking a health check to a DNS record for failover routing
A health check alone does not change DNS routing — it must be linked to a DNS record with a routing policy. For simple failover (primary/secondary), attach the health check ID to the primary record:
# Primary record — linked to health check
# Route 53 only returns this record when the health check is HEALTHY
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={
"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": "primary-us-east-1",
"Failover": "PRIMARY",
"TTL": 60,
"ResourceRecords": [
{"Value": "203.0.113.10"} # primary MCP server IP
],
# Attach health check — record is withheld when check is unhealthy
"HealthCheckId": health_check_id,
}
}]
}
)
# Secondary record — no health check required
# Returned only when the primary health check is unhealthy
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={
"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": "secondary-us-west-2",
"Failover": "SECONDARY",
"TTL": 60,
"ResourceRecords": [
{"Value": "203.0.113.20"} # secondary MCP server IP
],
}
}]
}
)
See the Route 53 failover routing guide for the full primary/secondary architecture including secondary health check linking and active-active configurations.
CloudWatch alarms and SNS notifications for health check failures
Route 53 publishes health check status as a CloudWatch metric in the us-east-1 region only (regardless of where your resources are). Create a CloudWatch alarm that triggers an SNS topic when the status drops to unhealthy:
import boto3
cloudwatch = boto3.client("cloudwatch", region_name="us-east-1")
sns = boto3.client("sns", region_name="us-east-1")
# Create SNS topic for health check alerts
topic_arn = sns.create_topic(Name="mcp-server-route53-alerts")["TopicArn"]
sns.subscribe(
TopicArn=topic_arn,
Protocol="email",
Endpoint="oncall@example.com",
)
# Create CloudWatch alarm on Route 53 HealthCheckStatus metric
# Metric: 1 = healthy, 0 = unhealthy
cloudwatch.put_metric_alarm(
AlarmName="mcp-server-route53-healthcheck",
MetricName="HealthCheckStatus",
Namespace="AWS/Route53",
Dimensions=[
{"Name": "HealthCheckId", "Value": health_check_id}
],
Statistic="Minimum",
Period=60, # evaluate every 60 seconds
EvaluationPeriods=1, # alarm after 1 consecutive unhealthy period
Threshold=1.0,
ComparisonOperator="LessThanThreshold",
AlarmActions=[topic_arn],
OKActions=[topic_arn],
TreatMissingData="breaching", # treat missing data as unhealthy
AlarmDescription=f"Route 53 health check failed for mcp.example.com (ID: {health_check_id})",
)
# Also enable Route 53's built-in SNS notification directly from the health check
route53.update_health_check(
HealthCheckId=health_check_id,
AlarmIdentifier={
"Region": "us-east-1", # always us-east-1 for Route 53 metrics
"Name": "mcp-server-route53-healthcheck",
},
InsufficientDataHealthStatus="LastKnownStatus",
)
Route 53 health check metrics are only published in us-east-1 — this is an AWS constraint that applies regardless of where the probed endpoint is hosted. Always create CloudWatch alarms for Route 53 health checks in us-east-1.
What Route 53 health checks do and do not verify for MCP servers
Understanding the gap between Route 53 health checks and full MCP protocol validation is essential for setting up a complete monitoring stack:
# What Route 53 health checks verify:
# ✓ TCP port is accepting connections
# ✓ HTTP response status is 2xx (for HTTP/HTTPS checks)
# ✓ TLS certificate is valid and not expired (for HTTPS checks)
# ✓ Response body contains a search string (optional, first 5120 bytes only)
# ✓ Response time is within probe timeout (10s default)
# What Route 53 health checks do NOT verify:
# ✗ JSON-RPC initialize request succeeds (MCP protocol handshake)
# ✗ tools/list returns a valid schema with expected tools
# ✗ Individual MCP tools execute correctly and return results
# ✗ Authentication (OAuth, API key) is accepted by the server
# ✗ Schema drift — tools/list response changed since last deployment
# ✗ Response latency at the MCP tool layer (P95/P99 for tool calls)
# ✗ Transport-specific issues (SSE reconnects, streamable HTTP behavior)
# The typical health endpoint only checks the HTTP layer:
@app.get("/health")
async def health():
return {"status": "ok"}
# This passes Route 53 health checks even when:
# - The MCP server's tool registry fails to load (tools/list returns empty)
# - The database is down (tools that query DB fail silently)
# - Auth middleware is misconfigured (all authenticated tool calls fail)
# - The server is in a half-initialized state after a crash-loop restart
Route 53 health checks are the right tool for DNS failover — they detect infrastructure-level failures (server unreachable, port closed, TLS expired) fast enough to update DNS. For MCP protocol-level validation — verifying that the server properly handles initialize, tools/list, and authenticated tool calls — use AliveMCP, which probes at the JSON-RPC layer and alerts on schema drift, tool-call failures, and authentication breakage that a simple HTTP health check misses entirely.
Calculated health checks for MCP server clusters
When an MCP server runs as a multi-instance cluster (multiple EC2 instances, ECS tasks, or App Runner instances behind a shared DNS name), a Calculated health check aggregates individual instance checks into a single cluster-level health signal:
# Create individual health checks for each instance
instance_checks = []
for i, instance_ip in enumerate(["203.0.113.10", "203.0.113.11", "203.0.113.12"]):
response = route53.create_health_check(
CallerReference=str(uuid.uuid4()),
HealthCheckConfig={
"Type": "HTTPS",
"IPAddress": instance_ip, # probe the specific instance IP
"FullyQualifiedDomainName": "mcp.example.com", # SNI hostname
"Port": 443,
"ResourcePath": "/health",
"RequestInterval": 30,
"FailureThreshold": 3,
"EnableSNI": True,
},
)
instance_checks.append(response["HealthCheck"]["Id"])
# Create a Calculated check: healthy if at least 2 of 3 instances are healthy
# This handles rolling deployments and single-instance failures
calculated = route53.create_health_check(
CallerReference=str(uuid.uuid4()),
HealthCheckConfig={
"Type": "CALCULATED",
"HealthThreshold": 2, # minimum healthy child checks required
"ChildHealthChecks": instance_checks,
# InvertHealthCheck: False (healthy when threshold met, unhealthy otherwise)
},
)
cluster_health_check_id = calculated["HealthCheck"]["Id"]
# Attach the Calculated check to the DNS record
# The DNS record stays healthy as long as 2 of 3 instances are healthy
# A single failed instance during a rolling deploy does not trigger failover
Go beyond HTTP pings with AliveMCP protocol monitoring
Route 53 health checks tell you when the server is unreachable. AliveMCP tells you when the MCP server is reachable but broken — failed tool calls, empty tool lists, schema drift after a deployment, and auth failures. Run both: Route 53 for DNS-layer failover, AliveMCP for protocol-layer alerting.
Join the waitlist →