Guide · AWS App Runner · MCP Server Deployment · Container Hosting
Deploy MCP Servers on AWS App Runner
AWS App Runner is the fastest path from a containerized MCP server to a production HTTPS endpoint — no VPCs, load balancers, clusters, or task definitions to configure. You point App Runner at an ECR image, set your port and environment variables, and it handles provisioning, TLS termination, health checks, and auto-scaling. For MCP server operators who want production-grade hosting without managing ECS clusters or Kubernetes, App Runner eliminates the infrastructure layer entirely. Three aspects matter most for MCP workloads: health check configuration (App Runner's defaults work poorly with long-lived SSE connections), IAM role separation (access role pulls the ECR image; instance role grants runtime AWS permissions to the MCP tools), and scale-to-zero trade-offs (MinSize=0 saves cost but introduces 10–30s cold starts on first MCP call after idle).
TL;DR
Create an App Runner service with apprunner.create_service(), pointing at an ECR private image. Set HealthCheckConfiguration.Protocol=HTTP and Path=/health with a long Timeout (10s) and high UnhealthyThreshold (10) to avoid App Runner evicting instances during active SSE streams. Use a dedicated instance role for MCP tool AWS credentials — App Runner injects it as the container's IAM identity. Set MinSize=1 for production to avoid cold starts.
Creating an App Runner service from an ECR image
The minimum viable App Runner service for a containerized MCP server:
import boto3
apprunner = boto3.client("apprunner", region_name="us-east-1")
response = apprunner.create_service(
ServiceName="mcp-server-prod",
SourceConfiguration={
"AuthenticationConfiguration": {
# IAM role that allows App Runner to pull from ECR
# Must trust principal: build.apprunner.amazonaws.com
"AccessRoleArn": "arn:aws:iam::123456789012:role/AppRunnerECRAccessRole"
},
"AutoDeploymentsEnabled": True, # re-deploy on new ECR image push
"ImageRepository": {
"ImageIdentifier": "123456789012.dkr.ecr.us-east-1.amazonaws.com/mcp-server:latest",
"ImageRepositoryType": "ECR",
"ImageConfiguration": {
"Port": "8080",
"RuntimeEnvironmentVariables": {
"LOG_LEVEL": "info",
"MCP_TRANSPORT": "sse",
"DATABASE_URL": "postgresql://...", # inject secrets here
},
},
},
},
InstanceConfiguration={
"Cpu": "1 vCPU",
"Memory": "2 GB",
# IAM role injected into the container as instance credentials
# MCP tool handlers use this to call AWS services
"InstanceRoleArn": "arn:aws:iam::123456789012:role/MCPServerInstanceRole",
},
HealthCheckConfiguration={
"Protocol": "HTTP",
"Path": "/health",
"Interval": 10, # seconds between health checks
"Timeout": 5, # seconds to wait for response
"HealthyThreshold": 1, # 1 success = healthy
"UnhealthyThreshold": 10, # 10 consecutive failures before eviction
},
)
The ServiceUrl in the response is the App Runner-managed HTTPS domain (format: {id}.{region}.awsapprunner.com). Custom domains can be added with associate_custom_domain().
IAM role architecture: access role vs instance role
App Runner uses two distinct IAM roles with different principals:
# Access role: App Runner pulls the ECR image on your behalf
# Trust policy principal: build.apprunner.amazonaws.com
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": {"Service": "build.apprunner.amazonaws.com"},
"Action": "sts:AssumeRole"
}]
}
# Attach managed policy: arn:aws:iam::aws:policy/service-role/AWSAppRunnerServicePolicyForECRAccess
# Instance role: credentials injected into the running container
# Trust policy principal: tasks.apprunner.amazonaws.com
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": {"Service": "tasks.apprunner.amazonaws.com"},
"Action": "sts:AssumeRole"
}]
}
# Attach your MCP server's permissions: DynamoDB, S3, Secrets Manager, etc.
Inside the container, the instance role is available as standard AWS credential chain — boto3.client("s3") will use it automatically via the container's metadata endpoint. No need to pass access keys in environment variables.
# In the MCP server container — instance role auto-resolves
import boto3
# These clients inherit the instance role credentials automatically
dynamodb = boto3.resource("dynamodb", region_name="us-east-1")
secrets = boto3.client("secretsmanager", region_name="us-east-1")
s3 = boto3.client("s3", region_name="us-east-1")
# For SSE-based MCP servers: fetch secrets at startup, cache in memory
def load_secrets():
response = secrets.get_secret_value(SecretId="mcp-server/prod/db-credentials")
import json
return json.loads(response["SecretString"])
Health check configuration for MCP servers with SSE
App Runner's default health check (Protocol=TCP) only checks that the port is accepting connections. For MCP servers, HTTP health checks are better — they verify the application stack is fully initialized. The critical parameter for SSE-based MCP servers is UnhealthyThreshold:
# Correct health check config for SSE-based MCP servers
HealthCheckConfiguration={
"Protocol": "HTTP",
"Path": "/health", # implement a lightweight health endpoint
"Interval": 10, # check every 10 seconds
"Timeout": 5, # 5s timeout per check
"HealthyThreshold": 1, # 1 success restores healthy state
"UnhealthyThreshold": 10, # 10 failures (100s) before eviction
}
# Why UnhealthyThreshold matters for SSE:
# App Runner health checks hit each instance periodically.
# During an active SSE stream, the health endpoint may be slow
# if the event loop is busy. A low UnhealthyThreshold (default=5)
# can cause App Runner to evict a healthy instance mid-stream.
# Setting it to 10 gives 100 seconds of grace before eviction.
The /health endpoint in the MCP server should be lightweight — it checks that the server started successfully and that any required database connections are alive, but should not block the event loop:
# FastAPI example — lightweight health check for App Runner
from fastapi import FastAPI
from fastapi.responses import JSONResponse
import asyncio
app = FastAPI()
_db_pool = None # set at startup
@app.get("/health")
async def health_check():
if _db_pool is None:
return JSONResponse({"status": "starting"}, status_code=503)
try:
# Lightweight ping — use timeout to avoid blocking health check
async with asyncio.timeout(2.0):
await _db_pool.execute("SELECT 1")
return {"status": "ok"}
except Exception as e:
return JSONResponse({"status": "error", "detail": str(e)}, status_code=503)
Port routing and HTTPS termination
App Runner terminates TLS at the load balancer layer and forwards plain HTTP to your container on the configured port. The container never sees HTTPS:
# App Runner routing:
# Client → HTTPS/443 → App Runner TLS termination → HTTP/{your_port} → container
# Container only needs to listen on plain HTTP
# Default port is 8080; set Port in ImageConfiguration to match
# For Node.js MCP servers (stdio → HTTP bridge):
import express from 'express';
const app = express();
app.listen(8080, '0.0.0.0'); # App Runner requires 0.0.0.0 binding, not 127.0.0.1
# For Python MCP servers:
import uvicorn
uvicorn.run(app, host="0.0.0.0", port=8080) # 0.0.0.0 is required
# App Runner forwards these headers from the TLS-terminated request:
# X-Forwarded-For: client IP
# X-Forwarded-Proto: https
# X-Forwarded-Port: 443
The container must bind to 0.0.0.0, not 127.0.0.1 (localhost). App Runner's load balancer forwards traffic from outside the container's loopback interface — a server bound only to 127.0.0.1 will fail health checks and never receive requests.
Instance sizing for MCP server workloads
App Runner offers a fixed menu of CPU/memory combinations. The right choice depends on how CPU-heavy the MCP tools are:
# App Runner instance size options (as of 2026):
# "0.25 vCPU" / "0.5 GB" — dev/test only; too small for concurrent MCP sessions
# "0.5 vCPU" / "1 GB" — lightweight tools with no ML inference
# "1 vCPU" / "2 GB" — standard MCP server (most deployments start here)
# "1 vCPU" / "3 GB" — MCP servers loading large in-memory indexes
# "1 vCPU" / "4 GB" — heavy in-memory state (vector stores, LLM tokenizers)
# "2 vCPU" / "4 GB" — CPU-intensive tools (PDF processing, image resizing)
# "4 vCPU" / "12 GB" — MCP servers running local ML inference
# Update instance size after service creation:
apprunner.update_service(
ServiceArn="arn:aws:apprunner:us-east-1:123456789012:service/mcp-server-prod/...",
InstanceConfiguration={
"Cpu": "2 vCPU",
"Memory": "4 GB",
"InstanceRoleArn": "arn:aws:iam::123456789012:role/MCPServerInstanceRole",
},
)
App Runner performs the instance size update as a rolling deployment — existing requests complete on the old instance size before the new instances take over. There is no downtime during a resize.
Checking service status and current deployment
App Runner services go through a state machine on creation and updates. Poll the service status to detect deployment failures before they affect traffic:
import time
def wait_for_service_running(service_arn: str, timeout_seconds: int = 600) -> dict:
"""Poll App Runner service until RUNNING or timeout."""
deadline = time.time() + timeout_seconds
while time.time() < deadline:
response = apprunner.describe_service(ServiceArn=service_arn)
service = response["Service"]
status = service["Status"]
if status == "RUNNING":
return service
if status in ("CREATE_FAILED", "DELETE_FAILED", "UPDATE_FAILED"):
raise RuntimeError(
f"Service reached terminal failure state: {status}. "
f"Check CloudWatch Logs at /aws/apprunner/{service['ServiceName']}"
)
print(f"Service status: {status} — waiting...")
time.sleep(15)
raise TimeoutError(f"Service did not reach RUNNING within {timeout_seconds}s")
# Service state transitions:
# OPERATION_IN_PROGRESS → RUNNING (successful create/update)
# OPERATION_IN_PROGRESS → CREATE_FAILED / UPDATE_FAILED (check system logs)
# RUNNING → PAUSED (after pause_service())
# PAUSED → RUNNING (after resume_service())
On CREATE_FAILED, inspect the system log group at /aws/apprunner/{service_name}/{service_id}/system — this contains the container pull errors, health check failures, and instance launch failures that caused the rollback.
App Runner vs ECS Fargate for MCP server deployments
Both services run containers on AWS-managed compute. The right choice depends on how much networking and operational control the MCP server requires:
- Choose App Runner when: the MCP server is stateless or uses managed databases (RDS, DynamoDB, ElastiCache) accessible via VPC connector; you want no-ops deployment; the team doesn't manage ECS clusters; cost savings from scale-to-zero matter for low-traffic tools.
- Choose ECS Fargate when: the MCP server needs fine-grained VPC subnet routing (e.g., placement in a specific AZ for latency); you need service mesh (App Mesh, ECS Service Connect); the MCP server talks to services that require specific security group rules App Runner's VPC connector can't express; you need sidecars (log shippers, service mesh proxies).
- Key App Runner limitation: each App Runner instance is a single container (no sidecar support). For MCP servers that need a sidecar agent (Datadog, Envoy proxy, log shipper), use ECS Fargate task definitions with multiple container definitions instead.
Monitor App Runner MCP endpoints with AliveMCP
App Runner services can reach unhealthy states silently — the service URL stays up but MCP tool calls return 502s or time out. AliveMCP probes each MCP server endpoint every 60 seconds, alerts on response failures, and tracks uptime history — so you know the moment your App Runner service degrades, before users do.
Join the waitlist →