Guide · AWS Route 53 · MCP Server High Availability · DNS Failover
Route 53 Failover Routing for MCP Server High Availability
Route 53 failover routing is the DNS-layer circuit breaker for MCP server deployments: when the primary endpoint fails a health check, Route 53 automatically stops returning it in DNS responses and starts returning the secondary endpoint — without any manual intervention. For MCP servers, this pattern matters because MCP clients resolve the server hostname at startup and may not retry DNS on every tool call. DNS-level failover is the only mechanism that redirects clients that cached a stale IP before the failure occurred. Three design decisions govern failover behavior: active-passive vs active-active (active-passive has one idle standby that absorbs traffic on failover; active-active shares load across both endpoints in normal operation), secondary failover target (a static error page, a maintenance-mode MCP server, or a full read-only replica depending on what users should experience during an outage), and DNS TTL on failover records (short TTLs reduce the propagation window for the failover switch but increase DNS query costs and resolver load during normal operation).
TL;DR
Create two Route 53 record sets with the same name and type but different Failover values: PRIMARY with a HealthCheckId attached, and SECONDARY without (or with a separate health check). Set TTL to 60s on both records. Route 53 returns the PRIMARY record while its health check passes; when the health check fails for FailureThreshold consecutive intervals, Route 53 switches to returning the SECONDARY record. Failback to PRIMARY happens automatically when the primary health check recovers.
Active-passive failover: primary and secondary MCP endpoints
The active-passive pattern keeps one endpoint live (primary) and one on standby (secondary). All traffic goes to the primary under normal conditions:
import boto3, uuid
route53 = boto3.client("route53")
# Step 1: Create health check for the primary MCP server endpoint
primary_hc = route53.create_health_check(
CallerReference=str(uuid.uuid4()),
HealthCheckConfig={
"Type": "HTTPS",
"FullyQualifiedDomainName": "primary.mcp-internal.example.com",
"Port": 443,
"ResourcePath": "/health",
"RequestInterval": 30,
"FailureThreshold": 3, # 90s to failover at 30s interval
"EnableSNI": True,
},
)
primary_hc_id = primary_hc["HealthCheck"]["Id"]
# Step 2: Create the PRIMARY record (returned when health check passes)
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={
"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
# SetIdentifier: unique within the same name+type
"SetIdentifier": "mcp-primary-us-east-1",
"Failover": "PRIMARY",
"TTL": 60,
"ResourceRecords": [{"Value": "203.0.113.10"}],
# Health check controls whether PRIMARY is included in responses
"HealthCheckId": primary_hc_id,
}
}]
}
)
# Step 3: Create the SECONDARY record (returned only when primary fails)
# Secondary does NOT require a health check — it acts as the last resort
# If the secondary also has a health check and it fails too,
# Route 53 returns the secondary anyway (to avoid a complete NXDOMAIN)
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={
"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": "mcp-secondary-us-west-2",
"Failover": "SECONDARY",
"TTL": 60,
"ResourceRecords": [{"Value": "203.0.113.20"}],
# Optional: attach a health check to the secondary too
# If secondary check also fails, Route 53 still returns secondary
# rather than returning NXDOMAIN — prevents total DNS blackout
}
}]
}
)
Failover timing: At 30s interval with FailureThreshold=3, Route 53 detects a failure after 90 seconds of continuous probe failures. After the check flips unhealthy, the DNS change propagates within the TTL — typically within 60–120 additional seconds (one TTL window for cached records to expire). Total failover window from failure to all-clients-redirected: approximately 3–4 minutes at default settings. Use RequestInterval=10 (fast health checks) to reduce detection to 30 seconds, at the cost of ~3× higher health check pricing.
Failover with Alias records for ALB targets
When the primary and secondary MCP servers are both behind Application Load Balancers, use Alias records instead of plain IP A records. Alias records with EvaluateTargetHealth: True layer ALB target group health on top of the Route 53 health check:
# Primary ALB Alias record with both Route 53 health check AND ALB evaluation
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={
"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": "mcp-primary-alb-us-east-1",
"Failover": "PRIMARY",
"AliasTarget": {
"DNSName": "mcp-primary-alb.us-east-1.elb.amazonaws.com",
"HostedZoneId": "Z35SXDOTRQ7X7K", # us-east-1 ALB hosted zone
# True = Route 53 considers the ALB unhealthy if all its
# registered targets are unhealthy (no additional health check needed)
"EvaluateTargetHealth": True,
},
# Attach an explicit Route 53 health check for /health path monitoring
# AND ALB target evaluation — both must pass for PRIMARY to be served
"HealthCheckId": primary_hc_id,
}
}]
}
)
# Secondary ALB Alias record — no explicit health check, EvaluateTargetHealth only
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={
"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": "mcp-secondary-alb-us-west-2",
"Failover": "SECONDARY",
"AliasTarget": {
"DNSName": "mcp-secondary-alb.us-west-2.elb.amazonaws.com",
"HostedZoneId": "Z1H1FL5HABSF5", # us-west-2 ALB hosted zone
"EvaluateTargetHealth": True,
},
}
}]
}
)
The dual-layer health check (Route 53 + ALB target evaluation) means Route 53 will fail over when any of these conditions occur: the health check path returns a non-2xx status, the health check probe times out, or the ALB has no healthy targets in its target groups. This provides strong guarantees without requiring separate health check resources for every target group behind the ALB.
Active-active failover with weighted routing
Active-active distributes traffic across multiple MCP endpoints normally, and removes a failed endpoint from rotation when its health check fails. Use weighted records (all with non-zero weights) and attach a health check to each:
# Active-active with weighted routing: 50/50 split in normal operation
# Route 53 removes an endpoint from rotation if its health check fails
# The remaining healthy endpoint(s) receive 100% of traffic during failures
endpoints = [
{"id": "us-east-1", "ip": "203.0.113.10", "weight": 50},
{"id": "us-west-2", "ip": "203.0.113.20", "weight": 50},
]
for endpoint in endpoints:
# Create a health check per endpoint
hc = route53.create_health_check(
CallerReference=str(uuid.uuid4()),
HealthCheckConfig={
"Type": "HTTPS",
"IPAddress": endpoint["ip"],
"FullyQualifiedDomainName": "mcp.example.com",
"Port": 443,
"ResourcePath": "/health",
"RequestInterval": 30,
"FailureThreshold": 3,
"EnableSNI": True,
},
)
# Create a weighted record for each endpoint
route53.change_resource_record_sets(
HostedZoneId=HOSTED_ZONE_ID,
ChangeBatch={
"Changes": [{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "mcp.example.com",
"Type": "A",
"SetIdentifier": f"mcp-active-{endpoint['id']}",
"Weight": endpoint["weight"],
"TTL": 60,
"ResourceRecords": [{"Value": endpoint["ip"]}],
"HealthCheckId": hc["HealthCheck"]["Id"],
}
}]
}
)
# With active-active weighted routing:
# Normal: resolver sees both records, randomly picks one (50/50 split)
# One failed: Route 53 omits the failed record → 100% to surviving endpoint
# Both failed: Route 53 returns BOTH records (to avoid NXDOMAIN) —
# clients will attempt to connect and fail; use active-passive
# if complete outage protection is more important than load distribution
Triggering manual failover for planned maintenance
Route 53 provides a way to force a health check into an unhealthy state without requiring an actual endpoint failure. This is the safest way to perform planned maintenance on the primary MCP server:
import time
def failover_to_secondary(health_check_id: str, reason: str):
"""Force the primary health check unhealthy to redirect traffic to secondary."""
route53.update_health_check(
HealthCheckId=health_check_id,
# Disabled = health check always evaluates to unhealthy
# This is equivalent to manually flipping the circuit breaker
Disabled=True,
)
print(f"Primary health check disabled: {reason}")
print("DNS will update within TTL window (~60s)")
def restore_primary(health_check_id: str):
"""Re-enable the primary health check to allow traffic back to primary."""
route53.update_health_check(
HealthCheckId=health_check_id,
Disabled=False,
)
print("Primary health check re-enabled. Monitoring for healthy status...")
# Poll until healthy
deadline = time.time() + 300 # wait up to 5 minutes
while time.time() < deadline:
response = route53.get_health_check_status(HealthCheckId=health_check_id)
statuses = [obs["StatusReport"]["Status"]
for obs in response["HealthCheckObservations"]]
healthy_count = sum(1 for s in statuses if s.startswith("Success"))
total = len(statuses)
print(f"Health check observers: {healthy_count}/{total} healthy")
if healthy_count >= total * 0.7: # 70% of probers healthy
print("Primary endpoint healthy — failback complete")
return
time.sleep(30)
# Planned maintenance workflow:
# 1. Disable primary health check → traffic moves to secondary
# 2. Perform maintenance on primary
# 3. Re-enable and monitor until primary is confirmed healthy
# 4. DNS automatically routes back to primary (no manual step needed)
Failover behavior edge cases for MCP deployments
Several edge cases affect how Route 53 failover behaves in practice:
- Active MCP sessions during failover: Route 53 DNS failover only affects new DNS resolutions. MCP clients that have already resolved the primary IP and established a session continue using that IP until the TCP connection drops or the session ends. DNS failover does not terminate existing sessions — the load balancer or server must close those connections for clients to re-connect and resolve the new IP. Design MCP server graceful shutdown to close SSE streams cleanly on SIGTERM so clients reconnect to the secondary immediately.
- Both health checks fail simultaneously: If both primary and secondary health checks fail, Route 53 returns both records anyway (it does not return NXDOMAIN). Clients will attempt connections and fail at the TCP or TLS layer. This is intentional — a complete DNS blackout is worse than serving unhealthy endpoints. For write-critical MCP servers, prefer an active-passive architecture where the secondary is a read-only maintenance page that returns clear JSON-RPC errors, rather than no response at all.
- Client DNS caching beyond TTL: Many MCP client frameworks (LLM agent runtimes, Claude Desktop, IDE extensions) cache the resolved IP for the process lifetime and do not re-resolve DNS on every request. A 60s TTL only affects clients that re-resolve. To guarantee failover for long-running agent processes, use load-balancer level health checks at the ALB target group layer — those route around failed targets without a DNS change.
- Failback timing: Route 53 restores the primary record as soon as the primary health check recovers (passes
HealthyThresholdconsecutive intervals, default 3). At 30s probe interval, failback takes a minimum of 90s from recovery. This means brief primary flaps can cause repeated failover-failback cycles (flapping). Use AliveMCP's minimum-stable-duration alert suppression alongside Route 53 failover to avoid alert fatigue during flapping events.
Detect MCP failover events with AliveMCP
Route 53 failover reroutes DNS — but it does not tell you whether the secondary MCP endpoint is serving correct JSON-RPC responses after taking over traffic. AliveMCP monitors both your primary and any secondary MCP endpoints continuously, alerting on protocol failures the moment they appear regardless of which DNS target is active.
Join the waitlist →