AWS App Runner · 2026-10-01 · App Runner arc
AWS App Runner for MCP Servers: Deployment Architecture, Auto-Scaling for SSE, and Production CI/CD
AWS App Runner is the fastest path from a containerized MCP server to a production HTTPS endpoint — no VPCs to configure, no ECS clusters to manage, no task definitions to maintain. You point App Runner at an ECR image, set a port and environment variables, and the platform handles TLS termination, health checks, auto-scaling, and rolling deployments. But App Runner has a set of non-obvious constraints that matter specifically for MCP server workloads: SSE connections change how auto-scaling works, the dual-role IAM model separates image-pull credentials from runtime permissions, and EgressType=VPC routes all traffic through your VPC — including public API calls — unless you account for it. This guide synthesizes five App Runner topics — service creation and basics, VPC connector for private resources, auto-scaling for SSE workloads, CloudWatch and X-Ray observability, and CI/CD with ECR and GitHub Actions — into three structural patterns that cover the full deployment lifecycle.
TL;DR
- IAM roles: App Runner uses two roles —
AccessRoleArn(trustbuild.apprunner.amazonaws.com) to pull the ECR image, andInstanceRoleArn(trusttasks.apprunner.amazonaws.com) for runtime AWS credentials inside the container. Never put runtime permissions in the access role. - SSE scaling: each MCP SSE session holds one open HTTP connection for its entire duration — set
MaxConcurrencyto your per-instance memory budget divided by memory-per-session (typically 10–50 for SSE, not the default 100). SetMinSize=1for production to avoid 15–50s cold starts. - VPC connector:
EgressType=VPCdisables public internet access from the service — add a NAT Gateway or VPC Interface Endpoints for any AWS service calls that would otherwise go over the public internet. - Observability: alarm on
5xxStatusResponses ≥ 1(60s window) andRequestLatency p99 > 5000ms(two 5-minute windows). Enable X-Ray viacreate_observability_configurationwithTraceConfiguration.Vendor=AWSXRAY; App Runner injects the daemon as a sidecar. - CI/CD: tag ECR images with the git SHA (
main-{sha}), not:latest. UseAutoDeploymentsEnabled=Falsein production and deploy explicitly viaupdate_service(). Rollback = re-runupdate_service()with the previous image URI.
Pattern 1 — Deployment and network architecture
Service creation: the minimal viable App Runner service
The core create_service() call for an SSE-based MCP server needs five things set correctly from the start — IAM roles, port configuration, health check parameters, container binding address, and instance sizing:
import boto3
apprunner = boto3.client("apprunner", region_name="us-east-1")
response = apprunner.create_service(
ServiceName="mcp-server-prod",
SourceConfiguration={
"AuthenticationConfiguration": {
# Pull role: trusts build.apprunner.amazonaws.com
# Policy: AWSAppRunnerServicePolicyForECRAccess (managed)
"AccessRoleArn": "arn:aws:iam::123456789012:role/AppRunnerECRAccessRole"
},
"AutoDeploymentsEnabled": False, # manual deploys for production
"ImageRepository": {
"ImageIdentifier": "123456789012.dkr.ecr.us-east-1.amazonaws.com/mcp-server:main-abc1234",
"ImageRepositoryType": "ECR",
"ImageConfiguration": {
"Port": "8080",
"RuntimeEnvironmentVariables": {
"LOG_LEVEL": "info",
"MCP_TRANSPORT": "sse",
},
},
},
},
InstanceConfiguration={
"Cpu": "1 vCPU",
"Memory": "2 GB",
# Runtime role: trusts tasks.apprunner.amazonaws.com
# Attach DynamoDB, S3, Secrets Manager permissions here
"InstanceRoleArn": "arn:aws:iam::123456789012:role/MCPServerInstanceRole",
},
HealthCheckConfiguration={
"Protocol": "HTTP",
"Path": "/health",
"Interval": 10,
"Timeout": 5,
"HealthyThreshold": 1,
"UnhealthyThreshold": 10, # 100s grace period — prevents eviction during SSE streams
},
)
Three details matter here. First, the IAM split: the AccessRoleArn is only for ECR image pulling and must trust the build.apprunner.amazonaws.com service principal — it never needs DynamoDB or S3 permissions. The InstanceRoleArn is what the container sees at runtime via the standard AWS credential chain — boto3.client("s3") picks it up automatically from the container metadata endpoint. Never put runtime permissions in the access role or inject access keys as environment variables.
Second, the container must bind to 0.0.0.0, not 127.0.0.1. App Runner's load balancer forwards requests from outside the container's loopback interface — a server bound only to localhost fails health checks silently and never serves requests.
Third, UnhealthyThreshold=10 (100 seconds of consecutive failures before eviction) is necessary for SSE servers. App Runner runs HTTP health checks on each instance while SSE sessions are active. If the event loop is busy processing a tool call, health check responses may be slow — the default threshold of 5 failures (50 seconds) can cause App Runner to evict a healthy instance mid-stream. Raising it to 10 provides enough grace to survive the busiest legitimate tool calls.
VPC connector: connecting to private RDS and ElastiCache
By default, App Runner services cannot reach VPC resources. The VPC connector creates an elastic network interface (ENI) inside your VPC subnets and routes all outbound traffic from the service through it:
# Step 1: Create the VPC connector (reusable across services)
connector = apprunner.create_vpc_connector(
VpcConnectorName="mcp-server-connector",
Subnets=[
"subnet-0abc123456789def0", # private subnet us-east-1a
"subnet-0def987654321abc0", # private subnet us-east-1b
],
SecurityGroups=["sg-0aaaaaaaaaaaaaaa1"],
)
connector_arn = connector["VpcConnector"]["VpcConnectorArn"]
# Step 2: Grant the connector's SG access to target resources
ec2 = boto3.client("ec2")
# Allow connector → RDS PostgreSQL
ec2.authorize_security_group_ingress(
GroupId="sg-0bbbbbbbbbbbbbb2", # RDS security group
IpPermissions=[{
"IpProtocol": "tcp", "FromPort": 5432, "ToPort": 5432,
"UserIdGroupPairs": [{"GroupId": "sg-0aaaaaaaaaaaaaaa1"}],
}],
)
# Allow connector → ElastiCache Redis
ec2.authorize_security_group_ingress(
GroupId="sg-0cccccccccccccc3", # ElastiCache security group
IpPermissions=[{
"IpProtocol": "tcp", "FromPort": 6379, "ToPort": 6379,
"UserIdGroupPairs": [{"GroupId": "sg-0aaaaaaaaaaaaaaa1"}],
}],
)
# Step 3: Attach the connector to the service
apprunner.update_service(
ServiceArn="arn:aws:apprunner:...",
NetworkConfiguration={
"EgressConfiguration": {
"EgressType": "VPC",
"VpcConnectorArn": connector_arn,
},
"IngressConfiguration": {"IsPubliclyAccessible": True},
},
)
A critical networking trap: EgressType=VPC routes all outbound traffic through the VPC, including calls to public AWS endpoints (S3, DynamoDB, Secrets Manager, Anthropic's Claude API). If the VPC subnets you attach don't have a route to a NAT Gateway — or VPC Interface Endpoints for the AWS services you use — those calls will fail silently. The connector's security group needs no inbound rules; it only handles egress traffic.
The connector itself costs $0.05/hr per connector-hour regardless of active connections. VPC connectors are immutable — to change subnets or security groups, create a new connector with the same name (this increments the version), update the service to reference the new connector ARN, then delete the old version. The old connector cannot be deleted while any service references it.
Inside the container, private RDS and ElastiCache hostnames resolve via the VPC's internal DNS resolver (the VPC CIDR +2 address, e.g., 10.0.0.2). The App Runner service resolves private RDS hostnames exactly as an EC2 instance or ECS task in the same VPC would — no special configuration in the application code.
Pattern 2 — Scaling and cost optimization
MaxConcurrency is a per-instance concurrent-connection threshold, not a request rate
App Runner's auto-scaling adds a new instance when any existing instance hits MaxConcurrency concurrent requests. For HTTP request/response workloads, a "concurrent request" lasts 50ms–5s — the counter rises and falls rapidly, and the default value of 100 is reasonable. For SSE-based MCP servers, each MCP session holds one open HTTP connection for its entire session duration — minutes to hours — and that connection counts as one concurrent request the whole time.
With the default MaxConcurrency=100, App Runner allows 100 simultaneous SSE sessions on a single 1 vCPU / 2 GB instance before adding another. At 80 MB per session (a realistic figure for a stateful MCP tool with a database connection pool and in-memory caches), 100 sessions would exhaust the 2 GB instance. The correct approach is to calculate MaxConcurrency from the memory budget:
# MaxConcurrency calculation for SSE-based MCP servers:
# instance_memory_gb = 2.0 # from InstanceConfiguration
# runtime_overhead_gb = 0.3 # OS + App Runner agent + Python/Node interpreter
# available_gb = 1.7 # 2.0 - 0.3
# memory_per_session = 80 # MB — measure via CloudWatch MemoryUtilization
# max_sessions = (1700 MB) / (80 MB) ≈ 21
# MaxConcurrency = 20 # slight buffer below the hard ceiling
# Create a custom auto-scaling configuration
asc = apprunner.create_auto_scaling_configuration(
AutoScalingConfigurationName="mcp-server-sse-scaling",
MaxConcurrency=20, # adds new instance when any instance hits 20 SSE sessions
MinSize=1, # always keep 1 instance running (no cold starts)
MaxSize=10, # cost cap: never exceed 10 instances
)
apprunner.update_service(
ServiceArn="arn:aws:apprunner:...",
AutoScalingConfigurationArn=asc["AutoScalingConfiguration"]["AutoScalingConfigurationArn"],
)
Watch the CloudWatch AWS/AppRunner Concurrency metric against MaxConcurrency × ActiveInstances. Sustained Concurrency near that product indicates the service is at capacity — scale-out will happen when the next session arrives. If P99 latency spikes correlate with instance count increases, lower MaxConcurrency to trigger earlier scale-out before individual instances become saturated.
Transport type also matters: HTTP + JSON-RPC MCP servers (short request/response cycles) can use MaxConcurrency=50–100 because connections are released after each tool call. SSE servers should use MaxConcurrency=10–50 because connections are held for the session lifetime. Lightweight stateless SSE tools (quick API calls, minimal in-memory state) can use 50; memory-heavy tools (embedding models, large in-memory indexes) should use 10–15.
Scale-to-zero vs MinSize=1: the cold start trade-off
Setting MinSize=0 allows App Runner to terminate all instances when there are no active connections, eliminating instance charges during idle periods. The cost is a 15–50s cold start on the first request after idle:
# Cold start timeline with MinSize=0:
# 1. First SSE connection arrives after idle period
# 2. App Runner detects zero instances → starts provisioning
# 3. ECR image pull: 5–15s (depends on image size; cached layers reduce this)
# 4. Container start: 5–15s (interpreter init + app startup)
# 5. Health check: Interval × HealthyThreshold = 10s × 1 = 10–20s (varies)
# 6. Request routed → total: 15–50s cold start
#
# MCP client sees: tool call hangs for 15–50 seconds on first use
# Often mistaken for a network error by users
# Minimize cold start time:
# - Keep the Docker image small (multi-stage builds, slim base images)
# - Defer heavy imports to first use (don't load ML models at module top-level)
# - Use lazy connection pools (create on first tool call, not at startup)
def get_heavy_client():
import anthropic # deferred — not loaded on cold start
return anthropic.Anthropic()
Use MinSize=0 for development, staging, and internal tools used during predictable business hours. Use MinSize=1 for production MCP servers where sub-second response time is required on the first session of the day. The cost difference is roughly one instance charge per month of idle time — for a 1 vCPU / 2 GB instance running 24/7, approximately $65/month. For production tools where a 30-second hang is unacceptable, this is worth paying.
Pause and resume for predictable off-hours savings
For production MCP servers with clear business-hours traffic patterns (e.g., an internal tool used 8 AM–6 PM on weekdays), the pause/resume pattern is more cost-effective than scale-to-zero. It eliminates cold starts during business hours while stopping instance charges during the predictable idle window:
import boto3
from datetime import datetime, timezone
apprunner = boto3.client("apprunner", region_name="us-east-1")
SERVICE_ARN = "arn:aws:apprunner:us-east-1:123456789012:service/mcp-server-prod/..."
def pause_service():
"""Stop all instances — hourly billing stops, service URL stays registered."""
apprunner.pause_service(ServiceArn=SERVICE_ARN)
def resume_service():
"""Restart instances — takes 30–90s to reach RUNNING."""
apprunner.resume_service(ServiceArn=SERVICE_ARN)
# Automate via EventBridge Scheduler:
# Pause: cron(0 22 ? * MON-FRI *) → 10 PM UTC weekdays
# Resume: cron(0 7 ? * MON-FRI *) → 7 AM UTC weekdays (1 hour before business start)
#
# 16 off-hours/day × 5 days/week × 4 weeks = 320 hours/month saved
# vs 720 total hours/month → ~44% cost reduction for a 5-day business-hours pattern
# vs 744 hours/month → higher savings on monthly comparison
During PAUSED state, inbound requests return 503 — the service URL stays registered and DNS doesn't change, so resumption is transparent. Resume time is 30–90s (image pull from ECR cache + container start + health check). Automating resume to fire 1 hour before business hours ensures the service is warm when the first user connects. This is strictly better than MinSize=0 for predictable patterns: scale-to-zero reacts to actual traffic, which means the first user of each session takes the cold start hit; pause/resume pre-warms on a schedule.
Pattern 3 — Observability and CI/CD
Two log groups and two alarm categories
App Runner creates two CloudWatch log groups automatically, with no agents or log shippers to configure:
# Application log group — everything written to stdout/stderr
# /aws/apprunner/{service_name}/{service_id}/application
# System log group — App Runner platform events
# /aws/apprunner/{service_name}/{service_id}/system
# Contains: deployment events, health check failures, instance replacements,
# scaling decisions, ECR pull failures → use this to diagnose
# UPDATE_FAILED and CREATE_FAILED states
# For useful CloudWatch Logs Insights queries, emit structured JSON to stdout:
import json, time, traceback
from contextvars import ContextVar
session_id: ContextVar[str] = ContextVar("session_id", default="unknown")
def log(level: str, tool: str, message: str, **extra):
entry = {
"timestamp": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
"level": level,
"tool": tool,
"message": message,
"session_id": session_id.get(),
**extra,
}
print(json.dumps(entry), flush=True) # flush=True for immediate delivery
# CloudWatch Logs Insights query: error rate by tool
# fields @timestamp, tool, error_type
# | filter level = "ERROR"
# | stats count(*) as error_count by tool, error_type
# | sort error_count desc
# P95 latency by tool:
# fields @timestamp, tool, duration_ms
# | filter level = "INFO" and message = "tool_call_success"
# | stats pct(duration_ms, 95) as p95_ms, count(*) as calls by tool
# | sort p95_ms desc
Two CloudWatch alarms cover the failure modes that matter most for MCP servers:
cloudwatch = boto3.client("cloudwatch", region_name="us-east-1")
SNS = "arn:aws:sns:us-east-1:123456789012:mcp-alerts"
# Alarm 1: any 5xx — broken MCP tools (surface immediately)
cloudwatch.put_metric_alarm(
AlarmName="mcp-server-5xx",
Namespace="AWS/AppRunner",
MetricName="5xxStatusResponses",
Dimensions=[{"Name": "ServiceName", "Value": "mcp-server-prod"}],
Statistic="Sum", Period=60, EvaluationPeriods=1, Threshold=1,
ComparisonOperator="GreaterThanOrEqualToThreshold",
TreatMissingData="notBreaching",
AlarmActions=[SNS], OKActions=[SNS],
)
# Alarm 2: P99 latency > 5s sustained — MCP tool degradation
cloudwatch.put_metric_alarm(
AlarmName="mcp-server-high-latency",
Namespace="AWS/AppRunner",
MetricName="RequestLatency",
Dimensions=[{"Name": "ServiceName", "Value": "mcp-server-prod"}],
ExtendedStatistic="p99", Period=300, EvaluationPeriods=2,
Threshold=5000,
ComparisonOperator="GreaterThanThreshold",
TreatMissingData="notBreaching",
AlarmActions=[SNS], OKActions=[SNS],
)
The 5xxStatusResponses ≥ 1 alarm uses a 60-second evaluation period and fires on a single occurrence. App Runner returns 5xx for several failure modes that matter for MCP tools: container returning 500 from an unhandled exception in a tool handler, container returning 503 during a deployment while health checks are still pending, or App Runner returning 502 when the VPC connector can't reach a backend resource. All of these are actionable — any 5xx in production warrants investigation.
The P99 latency alarm uses two 5-minute evaluation periods (10 minutes sustained) to avoid alerting on transient scale-out latency spikes, which are expected when a new instance starts up. Spikes that persist for more than two evaluation periods indicate a structural problem — backend database overload, a slow downstream API, or instances that are undersized for the current traffic pattern.
X-Ray tracing: connecting tool calls to downstream spans
X-Ray is enabled via an observability configuration rather than a Dockerfile change — App Runner injects the X-Ray daemon as a sidecar alongside the application container:
apprunner = boto3.client("apprunner")
# Step 1: Create the observability configuration
obs = apprunner.create_observability_configuration(
ObservabilityConfigurationName="mcp-server-xray",
TraceConfiguration={"Vendor": "AWSXRAY"},
)
# Step 2: Attach to the service
apprunner.update_service(
ServiceArn="arn:aws:apprunner:...",
ObservabilityConfiguration={
"ObservabilityEnabled": True,
"ObservabilityConfigurationArn": obs["ObservabilityConfiguration"]["ObservabilityConfigurationArn"],
},
)
# Step 3: In the container — instrument with aws_xray_sdk
from aws_xray_sdk.core import xray_recorder, patch_all
patch_all() # auto-instruments boto3, requests, aiohttp, httpx
async def my_mcp_tool(args: dict) -> str:
with xray_recorder.in_subsegment("mcp_tool_execution") as seg:
seg.put_annotation("tool_name", "fetch_workspace")
seg.put_annotation("session_id", session_id.get())
seg.put_metadata("args", args)
# boto3 calls inside appear as child spans automatically via patch_all()
result = await fetch_from_dynamodb(args["workspace_id"])
seg.put_metadata("result_size", len(str(result)))
return result
The daemon listens on 127.0.0.1:2000 UDP — no additional networking configuration is needed inside the container. patch_all() wraps every boto3 client in the process so DynamoDB queries, S3 reads, and Bedrock inference calls all appear as child spans under the parent App Runner request span. For MCP servers that call multiple downstream services per tool call, this trace view is the fastest way to identify which service is contributing latency.
CI/CD: immutable tags, explicit deploys, and rollback
The most important production CI/CD decision is tagging convention. Using :latest makes rollback impossible without knowing the specific image digest you need to revert to. Using git commit SHA tags makes every deployment auditable and every rollback a single update_service() call:
# GitHub Actions workflow for App Runner deployment
# .github/workflows/deploy.yml
name: Deploy MCP Server
on:
push:
branches: [main]
jobs:
deploy:
runs-on: ubuntu-latest
permissions:
id-token: write # OIDC
contents: read
steps:
- uses: actions/checkout@v4
- name: Configure AWS credentials (OIDC)
uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: arn:aws:iam::123456789012:role/GithubActionsDeployRole
aws-region: us-east-1
- name: Login to ECR
id: ecr
uses: aws-actions/amazon-ecr-login@v2
- name: Build and push
id: build
env:
REGISTRY: ${{ steps.ecr.outputs.registry }}
IMAGE_TAG: main-${{ github.sha }}
run: |
docker build -t $REGISTRY/mcp-server:$IMAGE_TAG .
docker push $REGISTRY/mcp-server:$IMAGE_TAG
echo "image_uri=$REGISTRY/mcp-server:$IMAGE_TAG" >> $GITHUB_OUTPUT
- name: Update App Runner service
run: |
aws apprunner update-service \
--service-arn ${{ secrets.APP_RUNNER_SERVICE_ARN }} \
--source-configuration "{
\"AuthenticationConfiguration\": {
\"AccessRoleArn\": \"${{ secrets.ECR_ACCESS_ROLE_ARN }}\"
},
\"AutoDeploymentsEnabled\": false,
\"ImageRepository\": {
\"ImageIdentifier\": \"${{ steps.build.outputs.image_uri }}\",
\"ImageRepositoryType\": \"ECR\",
\"ImageConfiguration\": {\"Port\": \"8080\"}
}
}"
- name: Wait for deployment
run: |
aws apprunner wait service-updated \
--service-arn ${{ secrets.APP_RUNNER_SERVICE_ARN }}
App Runner's update procedure is a managed rolling deployment: new instances start with the new image, health checks must pass before traffic shifts, and if health checks fail the rollback is automatic (old instances keep all traffic). Active SSE connections on old instances are maintained until the client disconnects or App Runner's drain timeout (~30–60s) expires. For MCP servers with long-lived sessions, this means some sessions may see the old version while new sessions see the new version during a deployment window — this is expected behavior.
Rollback when automatic rollback doesn't fire (e.g., the deployment succeeded but the new version has a subtle tool regression) is a single API call:
# Rollback to the previous image tag
apprunner.update_service(
ServiceArn="arn:aws:apprunner:...",
SourceConfiguration={
"AuthenticationConfiguration": {"AccessRoleArn": "arn:aws:iam::..."},
"AutoDeploymentsEnabled": False,
"ImageRepository": {
"ImageIdentifier": "123456789012.dkr.ecr.us-east-1.amazonaws.com/mcp-server:main-prevsha",
"ImageRepositoryType": "ECR",
"ImageConfiguration": {"Port": "8080"},
},
},
)
# Keep the last 10 ECR images (ECR lifecycle policy):
# {"countType": "imageCountMoreThan", "countNumber": 10,
# "selection": "taggedImages", "tagPrefixList": ["main-"]}
For secret rotation (e.g., rotating a database password in Secrets Manager), use start_deployment() to force a re-deploy of the current image without changing the service configuration. App Runner containers do not hot-reload environment variables — a start_deployment() triggers new containers that load the updated secret at startup while old containers drain gracefully.
Consolidated failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Service URL returns 502, health endpoint unreachable | Container bound to 127.0.0.1 instead of 0.0.0.0 |
Set host="0.0.0.0" in uvicorn/express listen call |
| Instances evicted mid-SSE-session, clients disconnected | UnhealthyThreshold too low (default 5); health check fails when event loop is busy |
Set UnhealthyThreshold=10 (100s grace) in HealthCheckConfiguration |
| RDS/ElastiCache connections time out with VPC connector attached | Connector SG not added as inbound source on RDS/ElastiCache SG | Add authorize_security_group_ingress rule with connector SG as source on port 5432/6379 |
| Public AWS API calls fail (S3, Secrets Manager) after VPC connector attach | EgressType=VPC routes all traffic through VPC; subnets lack NAT Gateway or VPC Interface Endpoints |
Add NAT Gateway to App Runner subnets, or create VPC Interface Endpoints for required services |
| Memory OOM on instances with multiple concurrent SSE sessions | MaxConcurrency too high for SSE workload; each session holds memory for its duration |
Set MaxConcurrency = (instance_memory_MB − 300) / memory_per_session_MB |
| 15–50s hang on first tool call after idle | MinSize=0 — service scaled to zero, cold start on first request |
Set MinSize=1 for production, or use pause/resume schedule to pre-warm before business hours |
CREATE_FAILED / UPDATE_FAILED on deployment |
Container fails health checks — application startup error, wrong port, missing env var | Check /aws/apprunner/{name}/{id}/system log group for deployment event details |
boto3 calls fail with NoCredentialsError inside container |
InstanceRoleArn not attached, or role trust policy uses wrong principal |
Trust principal must be tasks.apprunner.amazonaws.com; attach permissions in instance role (not access role) |
| Rollback to previous version impossible after bad deploy | :latest ECR tag was used — previous image hash unknown |
Always tag production images with git SHA (main-{sha}); keep last 10 in ECR lifecycle policy |
| Rotated secret not picked up by running container | App Runner containers do not hot-reload env vars; old containers keep cached secret value | Call start_deployment() after rotating — forces new containers that load the updated value at startup |
| X-Ray traces show no downstream spans for boto3 calls | patch_all() not called, or called after boto3 clients were already instantiated |
Call patch_all() at application start before any boto3 client creation |
| VPC connector update blocked — can't delete old connector version | Service still references the old connector ARN | Update service to reference new connector ARN first; wait for RUNNING; then delete old connector |
Production checklists
Service creation checklist
- Access role trusts
build.apprunner.amazonaws.com; hasAWSAppRunnerServicePolicyForECRAccessmanaged policy attached - Instance role trusts
tasks.apprunner.amazonaws.com; has runtime permissions (DynamoDB, S3, Secrets Manager, etc.) - Container binds to
0.0.0.0, not127.0.0.1 HealthCheckConfiguration.Protocol=HTTP,Path=/health,UnhealthyThreshold=10- Health endpoint returns 200 quickly without blocking the event loop; returns 503 until the DB pool is initialized
MinSize=1for production;MinSize=0only for dev/staging- Custom auto-scaling configuration with
MaxConcurrencyderived from memory budget per session
VPC connector checklist
- Connector placed in private subnets with routes to RDS/ElastiCache subnet groups
- Connector SG added as inbound source to RDS SG (port 5432) and ElastiCache SG (port 6379)
- Connector SG has no inbound rules (only handles egress)
- NAT Gateway or VPC Interface Endpoints in connector subnets for required AWS service calls
- VPC Interface Endpoints for: Secrets Manager, S3, DynamoDB, Bedrock — whichever the MCP tools call
- Old connector versions deleted after service update completes ($0.05/hr per idle connector)
Observability checklist
- Structured JSON logging to stdout with
flush=True; fields: timestamp, level, tool, session_id, duration_ms - CloudWatch alarm:
5xxStatusResponses ≥ 1in 60s window → SNS alert - CloudWatch alarm:
RequestLatency p99 > 5000msfor two 5-minute windows → SNS alert - CloudWatch alarm:
ActiveInstances > {expected_max}→ SNS alert for unexpected scaling - Observability configuration created with
TraceConfiguration.Vendor=AWSXRAYand attached to service patch_all()called at application start before any boto3 client instantiation
CI/CD checklist
- Production images tagged with git SHA (
main-{sha}), never:latest AutoDeploymentsEnabled=Falseon production service; deploy explicitly via CI pipeline- CI pipeline uses OIDC-based AWS auth (not long-lived access keys)
- CI waits for
aws apprunner wait service-updatedbefore marking deployment success - ECR lifecycle policy retains last 10 tagged images for rollback
- Rollback procedure documented:
update_service()with previous image URI - Secret rotation procedure documented: rotate in Secrets Manager → call
start_deployment()
Monitor your App Runner MCP endpoints end-to-end
App Runner's internal observability covers what happens inside your service — but it misses the full network path from the internet. AliveMCP probes each MCP endpoint every 60 seconds from outside AWS, measuring end-to-end response time including DNS resolution, TLS handshake, and the complete path through App Runner's load balancer. CloudWatch metrics stay quiet during DNS propagation failures and App Runner load balancer degradation — external probing catches those before your users do.
Join the waitlist →