AWS App Runner · 2026-10-01 · App Runner arc

AWS App Runner for MCP Servers: Deployment Architecture, Auto-Scaling for SSE, and Production CI/CD

AWS App Runner is the fastest path from a containerized MCP server to a production HTTPS endpoint — no VPCs to configure, no ECS clusters to manage, no task definitions to maintain. You point App Runner at an ECR image, set a port and environment variables, and the platform handles TLS termination, health checks, auto-scaling, and rolling deployments. But App Runner has a set of non-obvious constraints that matter specifically for MCP server workloads: SSE connections change how auto-scaling works, the dual-role IAM model separates image-pull credentials from runtime permissions, and EgressType=VPC routes all traffic through your VPC — including public API calls — unless you account for it. This guide synthesizes five App Runner topics — service creation and basics, VPC connector for private resources, auto-scaling for SSE workloads, CloudWatch and X-Ray observability, and CI/CD with ECR and GitHub Actions — into three structural patterns that cover the full deployment lifecycle.

TL;DR

Pattern 1 — Deployment and network architecture

Service creation: the minimal viable App Runner service

The core create_service() call for an SSE-based MCP server needs five things set correctly from the start — IAM roles, port configuration, health check parameters, container binding address, and instance sizing:

import boto3

apprunner = boto3.client("apprunner", region_name="us-east-1")

response = apprunner.create_service(
    ServiceName="mcp-server-prod",
    SourceConfiguration={
        "AuthenticationConfiguration": {
            # Pull role: trusts build.apprunner.amazonaws.com
            # Policy: AWSAppRunnerServicePolicyForECRAccess (managed)
            "AccessRoleArn": "arn:aws:iam::123456789012:role/AppRunnerECRAccessRole"
        },
        "AutoDeploymentsEnabled": False,  # manual deploys for production
        "ImageRepository": {
            "ImageIdentifier": "123456789012.dkr.ecr.us-east-1.amazonaws.com/mcp-server:main-abc1234",
            "ImageRepositoryType": "ECR",
            "ImageConfiguration": {
                "Port": "8080",
                "RuntimeEnvironmentVariables": {
                    "LOG_LEVEL": "info",
                    "MCP_TRANSPORT": "sse",
                },
            },
        },
    },
    InstanceConfiguration={
        "Cpu": "1 vCPU",
        "Memory": "2 GB",
        # Runtime role: trusts tasks.apprunner.amazonaws.com
        # Attach DynamoDB, S3, Secrets Manager permissions here
        "InstanceRoleArn": "arn:aws:iam::123456789012:role/MCPServerInstanceRole",
    },
    HealthCheckConfiguration={
        "Protocol": "HTTP",
        "Path": "/health",
        "Interval": 10,
        "Timeout": 5,
        "HealthyThreshold": 1,
        "UnhealthyThreshold": 10,  # 100s grace period — prevents eviction during SSE streams
    },
)

Three details matter here. First, the IAM split: the AccessRoleArn is only for ECR image pulling and must trust the build.apprunner.amazonaws.com service principal — it never needs DynamoDB or S3 permissions. The InstanceRoleArn is what the container sees at runtime via the standard AWS credential chain — boto3.client("s3") picks it up automatically from the container metadata endpoint. Never put runtime permissions in the access role or inject access keys as environment variables.

Second, the container must bind to 0.0.0.0, not 127.0.0.1. App Runner's load balancer forwards requests from outside the container's loopback interface — a server bound only to localhost fails health checks silently and never serves requests.

Third, UnhealthyThreshold=10 (100 seconds of consecutive failures before eviction) is necessary for SSE servers. App Runner runs HTTP health checks on each instance while SSE sessions are active. If the event loop is busy processing a tool call, health check responses may be slow — the default threshold of 5 failures (50 seconds) can cause App Runner to evict a healthy instance mid-stream. Raising it to 10 provides enough grace to survive the busiest legitimate tool calls.

VPC connector: connecting to private RDS and ElastiCache

By default, App Runner services cannot reach VPC resources. The VPC connector creates an elastic network interface (ENI) inside your VPC subnets and routes all outbound traffic from the service through it:

# Step 1: Create the VPC connector (reusable across services)
connector = apprunner.create_vpc_connector(
    VpcConnectorName="mcp-server-connector",
    Subnets=[
        "subnet-0abc123456789def0",  # private subnet us-east-1a
        "subnet-0def987654321abc0",  # private subnet us-east-1b
    ],
    SecurityGroups=["sg-0aaaaaaaaaaaaaaa1"],
)
connector_arn = connector["VpcConnector"]["VpcConnectorArn"]

# Step 2: Grant the connector's SG access to target resources
ec2 = boto3.client("ec2")
# Allow connector → RDS PostgreSQL
ec2.authorize_security_group_ingress(
    GroupId="sg-0bbbbbbbbbbbbbb2",  # RDS security group
    IpPermissions=[{
        "IpProtocol": "tcp", "FromPort": 5432, "ToPort": 5432,
        "UserIdGroupPairs": [{"GroupId": "sg-0aaaaaaaaaaaaaaa1"}],
    }],
)
# Allow connector → ElastiCache Redis
ec2.authorize_security_group_ingress(
    GroupId="sg-0cccccccccccccc3",  # ElastiCache security group
    IpPermissions=[{
        "IpProtocol": "tcp", "FromPort": 6379, "ToPort": 6379,
        "UserIdGroupPairs": [{"GroupId": "sg-0aaaaaaaaaaaaaaa1"}],
    }],
)

# Step 3: Attach the connector to the service
apprunner.update_service(
    ServiceArn="arn:aws:apprunner:...",
    NetworkConfiguration={
        "EgressConfiguration": {
            "EgressType": "VPC",
            "VpcConnectorArn": connector_arn,
        },
        "IngressConfiguration": {"IsPubliclyAccessible": True},
    },
)

A critical networking trap: EgressType=VPC routes all outbound traffic through the VPC, including calls to public AWS endpoints (S3, DynamoDB, Secrets Manager, Anthropic's Claude API). If the VPC subnets you attach don't have a route to a NAT Gateway — or VPC Interface Endpoints for the AWS services you use — those calls will fail silently. The connector's security group needs no inbound rules; it only handles egress traffic.

The connector itself costs $0.05/hr per connector-hour regardless of active connections. VPC connectors are immutable — to change subnets or security groups, create a new connector with the same name (this increments the version), update the service to reference the new connector ARN, then delete the old version. The old connector cannot be deleted while any service references it.

Inside the container, private RDS and ElastiCache hostnames resolve via the VPC's internal DNS resolver (the VPC CIDR +2 address, e.g., 10.0.0.2). The App Runner service resolves private RDS hostnames exactly as an EC2 instance or ECS task in the same VPC would — no special configuration in the application code.

Pattern 2 — Scaling and cost optimization

MaxConcurrency is a per-instance concurrent-connection threshold, not a request rate

App Runner's auto-scaling adds a new instance when any existing instance hits MaxConcurrency concurrent requests. For HTTP request/response workloads, a "concurrent request" lasts 50ms–5s — the counter rises and falls rapidly, and the default value of 100 is reasonable. For SSE-based MCP servers, each MCP session holds one open HTTP connection for its entire session duration — minutes to hours — and that connection counts as one concurrent request the whole time.

With the default MaxConcurrency=100, App Runner allows 100 simultaneous SSE sessions on a single 1 vCPU / 2 GB instance before adding another. At 80 MB per session (a realistic figure for a stateful MCP tool with a database connection pool and in-memory caches), 100 sessions would exhaust the 2 GB instance. The correct approach is to calculate MaxConcurrency from the memory budget:

# MaxConcurrency calculation for SSE-based MCP servers:
#   instance_memory_gb  = 2.0   # from InstanceConfiguration
#   runtime_overhead_gb = 0.3   # OS + App Runner agent + Python/Node interpreter
#   available_gb        = 1.7   # 2.0 - 0.3
#   memory_per_session  = 80    # MB — measure via CloudWatch MemoryUtilization
#   max_sessions        = (1700 MB) / (80 MB) ≈ 21
#   MaxConcurrency      = 20    # slight buffer below the hard ceiling

# Create a custom auto-scaling configuration
asc = apprunner.create_auto_scaling_configuration(
    AutoScalingConfigurationName="mcp-server-sse-scaling",
    MaxConcurrency=20,   # adds new instance when any instance hits 20 SSE sessions
    MinSize=1,           # always keep 1 instance running (no cold starts)
    MaxSize=10,          # cost cap: never exceed 10 instances
)

apprunner.update_service(
    ServiceArn="arn:aws:apprunner:...",
    AutoScalingConfigurationArn=asc["AutoScalingConfiguration"]["AutoScalingConfigurationArn"],
)

Watch the CloudWatch AWS/AppRunner Concurrency metric against MaxConcurrency × ActiveInstances. Sustained Concurrency near that product indicates the service is at capacity — scale-out will happen when the next session arrives. If P99 latency spikes correlate with instance count increases, lower MaxConcurrency to trigger earlier scale-out before individual instances become saturated.

Transport type also matters: HTTP + JSON-RPC MCP servers (short request/response cycles) can use MaxConcurrency=50–100 because connections are released after each tool call. SSE servers should use MaxConcurrency=10–50 because connections are held for the session lifetime. Lightweight stateless SSE tools (quick API calls, minimal in-memory state) can use 50; memory-heavy tools (embedding models, large in-memory indexes) should use 10–15.

Scale-to-zero vs MinSize=1: the cold start trade-off

Setting MinSize=0 allows App Runner to terminate all instances when there are no active connections, eliminating instance charges during idle periods. The cost is a 15–50s cold start on the first request after idle:

# Cold start timeline with MinSize=0:
# 1. First SSE connection arrives after idle period
# 2. App Runner detects zero instances → starts provisioning
# 3. ECR image pull: 5–15s (depends on image size; cached layers reduce this)
# 4. Container start: 5–15s (interpreter init + app startup)
# 5. Health check: Interval × HealthyThreshold = 10s × 1 = 10–20s (varies)
# 6. Request routed → total: 15–50s cold start
#
# MCP client sees: tool call hangs for 15–50 seconds on first use
# Often mistaken for a network error by users

# Minimize cold start time:
# - Keep the Docker image small (multi-stage builds, slim base images)
# - Defer heavy imports to first use (don't load ML models at module top-level)
# - Use lazy connection pools (create on first tool call, not at startup)
def get_heavy_client():
    import anthropic  # deferred — not loaded on cold start
    return anthropic.Anthropic()

Use MinSize=0 for development, staging, and internal tools used during predictable business hours. Use MinSize=1 for production MCP servers where sub-second response time is required on the first session of the day. The cost difference is roughly one instance charge per month of idle time — for a 1 vCPU / 2 GB instance running 24/7, approximately $65/month. For production tools where a 30-second hang is unacceptable, this is worth paying.

Pause and resume for predictable off-hours savings

For production MCP servers with clear business-hours traffic patterns (e.g., an internal tool used 8 AM–6 PM on weekdays), the pause/resume pattern is more cost-effective than scale-to-zero. It eliminates cold starts during business hours while stopping instance charges during the predictable idle window:

import boto3
from datetime import datetime, timezone

apprunner = boto3.client("apprunner", region_name="us-east-1")
SERVICE_ARN = "arn:aws:apprunner:us-east-1:123456789012:service/mcp-server-prod/..."

def pause_service():
    """Stop all instances — hourly billing stops, service URL stays registered."""
    apprunner.pause_service(ServiceArn=SERVICE_ARN)

def resume_service():
    """Restart instances — takes 30–90s to reach RUNNING."""
    apprunner.resume_service(ServiceArn=SERVICE_ARN)

# Automate via EventBridge Scheduler:
# Pause: cron(0 22 ? * MON-FRI *)  → 10 PM UTC weekdays
# Resume: cron(0 7 ? * MON-FRI *)  → 7 AM UTC weekdays (1 hour before business start)
#
# 16 off-hours/day × 5 days/week × 4 weeks = 320 hours/month saved
# vs 720 total hours/month → ~44% cost reduction for a 5-day business-hours pattern
# vs 744 hours/month → higher savings on monthly comparison

During PAUSED state, inbound requests return 503 — the service URL stays registered and DNS doesn't change, so resumption is transparent. Resume time is 30–90s (image pull from ECR cache + container start + health check). Automating resume to fire 1 hour before business hours ensures the service is warm when the first user connects. This is strictly better than MinSize=0 for predictable patterns: scale-to-zero reacts to actual traffic, which means the first user of each session takes the cold start hit; pause/resume pre-warms on a schedule.

Pattern 3 — Observability and CI/CD

Two log groups and two alarm categories

App Runner creates two CloudWatch log groups automatically, with no agents or log shippers to configure:

# Application log group — everything written to stdout/stderr
# /aws/apprunner/{service_name}/{service_id}/application

# System log group — App Runner platform events
# /aws/apprunner/{service_name}/{service_id}/system
# Contains: deployment events, health check failures, instance replacements,
#           scaling decisions, ECR pull failures → use this to diagnose
#           UPDATE_FAILED and CREATE_FAILED states

# For useful CloudWatch Logs Insights queries, emit structured JSON to stdout:
import json, time, traceback
from contextvars import ContextVar

session_id: ContextVar[str] = ContextVar("session_id", default="unknown")

def log(level: str, tool: str, message: str, **extra):
    entry = {
        "timestamp": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "level": level,
        "tool": tool,
        "message": message,
        "session_id": session_id.get(),
        **extra,
    }
    print(json.dumps(entry), flush=True)  # flush=True for immediate delivery

# CloudWatch Logs Insights query: error rate by tool
# fields @timestamp, tool, error_type
# | filter level = "ERROR"
# | stats count(*) as error_count by tool, error_type
# | sort error_count desc

# P95 latency by tool:
# fields @timestamp, tool, duration_ms
# | filter level = "INFO" and message = "tool_call_success"
# | stats pct(duration_ms, 95) as p95_ms, count(*) as calls by tool
# | sort p95_ms desc

Two CloudWatch alarms cover the failure modes that matter most for MCP servers:

cloudwatch = boto3.client("cloudwatch", region_name="us-east-1")
SNS = "arn:aws:sns:us-east-1:123456789012:mcp-alerts"

# Alarm 1: any 5xx — broken MCP tools (surface immediately)
cloudwatch.put_metric_alarm(
    AlarmName="mcp-server-5xx",
    Namespace="AWS/AppRunner",
    MetricName="5xxStatusResponses",
    Dimensions=[{"Name": "ServiceName", "Value": "mcp-server-prod"}],
    Statistic="Sum", Period=60, EvaluationPeriods=1, Threshold=1,
    ComparisonOperator="GreaterThanOrEqualToThreshold",
    TreatMissingData="notBreaching",
    AlarmActions=[SNS], OKActions=[SNS],
)

# Alarm 2: P99 latency > 5s sustained — MCP tool degradation
cloudwatch.put_metric_alarm(
    AlarmName="mcp-server-high-latency",
    Namespace="AWS/AppRunner",
    MetricName="RequestLatency",
    Dimensions=[{"Name": "ServiceName", "Value": "mcp-server-prod"}],
    ExtendedStatistic="p99", Period=300, EvaluationPeriods=2,
    Threshold=5000,
    ComparisonOperator="GreaterThanThreshold",
    TreatMissingData="notBreaching",
    AlarmActions=[SNS], OKActions=[SNS],
)

The 5xxStatusResponses ≥ 1 alarm uses a 60-second evaluation period and fires on a single occurrence. App Runner returns 5xx for several failure modes that matter for MCP tools: container returning 500 from an unhandled exception in a tool handler, container returning 503 during a deployment while health checks are still pending, or App Runner returning 502 when the VPC connector can't reach a backend resource. All of these are actionable — any 5xx in production warrants investigation.

The P99 latency alarm uses two 5-minute evaluation periods (10 minutes sustained) to avoid alerting on transient scale-out latency spikes, which are expected when a new instance starts up. Spikes that persist for more than two evaluation periods indicate a structural problem — backend database overload, a slow downstream API, or instances that are undersized for the current traffic pattern.

X-Ray tracing: connecting tool calls to downstream spans

X-Ray is enabled via an observability configuration rather than a Dockerfile change — App Runner injects the X-Ray daemon as a sidecar alongside the application container:

apprunner = boto3.client("apprunner")

# Step 1: Create the observability configuration
obs = apprunner.create_observability_configuration(
    ObservabilityConfigurationName="mcp-server-xray",
    TraceConfiguration={"Vendor": "AWSXRAY"},
)

# Step 2: Attach to the service
apprunner.update_service(
    ServiceArn="arn:aws:apprunner:...",
    ObservabilityConfiguration={
        "ObservabilityEnabled": True,
        "ObservabilityConfigurationArn": obs["ObservabilityConfiguration"]["ObservabilityConfigurationArn"],
    },
)

# Step 3: In the container — instrument with aws_xray_sdk
from aws_xray_sdk.core import xray_recorder, patch_all

patch_all()  # auto-instruments boto3, requests, aiohttp, httpx

async def my_mcp_tool(args: dict) -> str:
    with xray_recorder.in_subsegment("mcp_tool_execution") as seg:
        seg.put_annotation("tool_name", "fetch_workspace")
        seg.put_annotation("session_id", session_id.get())
        seg.put_metadata("args", args)
        # boto3 calls inside appear as child spans automatically via patch_all()
        result = await fetch_from_dynamodb(args["workspace_id"])
        seg.put_metadata("result_size", len(str(result)))
        return result

The daemon listens on 127.0.0.1:2000 UDP — no additional networking configuration is needed inside the container. patch_all() wraps every boto3 client in the process so DynamoDB queries, S3 reads, and Bedrock inference calls all appear as child spans under the parent App Runner request span. For MCP servers that call multiple downstream services per tool call, this trace view is the fastest way to identify which service is contributing latency.

CI/CD: immutable tags, explicit deploys, and rollback

The most important production CI/CD decision is tagging convention. Using :latest makes rollback impossible without knowing the specific image digest you need to revert to. Using git commit SHA tags makes every deployment auditable and every rollback a single update_service() call:

# GitHub Actions workflow for App Runner deployment
# .github/workflows/deploy.yml

name: Deploy MCP Server
on:
  push:
    branches: [main]

jobs:
  deploy:
    runs-on: ubuntu-latest
    permissions:
      id-token: write  # OIDC
      contents: read

    steps:
      - uses: actions/checkout@v4

      - name: Configure AWS credentials (OIDC)
        uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::123456789012:role/GithubActionsDeployRole
          aws-region: us-east-1

      - name: Login to ECR
        id: ecr
        uses: aws-actions/amazon-ecr-login@v2

      - name: Build and push
        id: build
        env:
          REGISTRY: ${{ steps.ecr.outputs.registry }}
          IMAGE_TAG: main-${{ github.sha }}
        run: |
          docker build -t $REGISTRY/mcp-server:$IMAGE_TAG .
          docker push $REGISTRY/mcp-server:$IMAGE_TAG
          echo "image_uri=$REGISTRY/mcp-server:$IMAGE_TAG" >> $GITHUB_OUTPUT

      - name: Update App Runner service
        run: |
          aws apprunner update-service \
            --service-arn ${{ secrets.APP_RUNNER_SERVICE_ARN }} \
            --source-configuration "{
              \"AuthenticationConfiguration\": {
                \"AccessRoleArn\": \"${{ secrets.ECR_ACCESS_ROLE_ARN }}\"
              },
              \"AutoDeploymentsEnabled\": false,
              \"ImageRepository\": {
                \"ImageIdentifier\": \"${{ steps.build.outputs.image_uri }}\",
                \"ImageRepositoryType\": \"ECR\",
                \"ImageConfiguration\": {\"Port\": \"8080\"}
              }
            }"

      - name: Wait for deployment
        run: |
          aws apprunner wait service-updated \
            --service-arn ${{ secrets.APP_RUNNER_SERVICE_ARN }}

App Runner's update procedure is a managed rolling deployment: new instances start with the new image, health checks must pass before traffic shifts, and if health checks fail the rollback is automatic (old instances keep all traffic). Active SSE connections on old instances are maintained until the client disconnects or App Runner's drain timeout (~30–60s) expires. For MCP servers with long-lived sessions, this means some sessions may see the old version while new sessions see the new version during a deployment window — this is expected behavior.

Rollback when automatic rollback doesn't fire (e.g., the deployment succeeded but the new version has a subtle tool regression) is a single API call:

# Rollback to the previous image tag
apprunner.update_service(
    ServiceArn="arn:aws:apprunner:...",
    SourceConfiguration={
        "AuthenticationConfiguration": {"AccessRoleArn": "arn:aws:iam::..."},
        "AutoDeploymentsEnabled": False,
        "ImageRepository": {
            "ImageIdentifier": "123456789012.dkr.ecr.us-east-1.amazonaws.com/mcp-server:main-prevsha",
            "ImageRepositoryType": "ECR",
            "ImageConfiguration": {"Port": "8080"},
        },
    },
)
# Keep the last 10 ECR images (ECR lifecycle policy):
# {"countType": "imageCountMoreThan", "countNumber": 10,
#  "selection": "taggedImages", "tagPrefixList": ["main-"]}

For secret rotation (e.g., rotating a database password in Secrets Manager), use start_deployment() to force a re-deploy of the current image without changing the service configuration. App Runner containers do not hot-reload environment variables — a start_deployment() triggers new containers that load the updated secret at startup while old containers drain gracefully.

Consolidated failure modes

Symptom Cause Fix
Service URL returns 502, health endpoint unreachable Container bound to 127.0.0.1 instead of 0.0.0.0 Set host="0.0.0.0" in uvicorn/express listen call
Instances evicted mid-SSE-session, clients disconnected UnhealthyThreshold too low (default 5); health check fails when event loop is busy Set UnhealthyThreshold=10 (100s grace) in HealthCheckConfiguration
RDS/ElastiCache connections time out with VPC connector attached Connector SG not added as inbound source on RDS/ElastiCache SG Add authorize_security_group_ingress rule with connector SG as source on port 5432/6379
Public AWS API calls fail (S3, Secrets Manager) after VPC connector attach EgressType=VPC routes all traffic through VPC; subnets lack NAT Gateway or VPC Interface Endpoints Add NAT Gateway to App Runner subnets, or create VPC Interface Endpoints for required services
Memory OOM on instances with multiple concurrent SSE sessions MaxConcurrency too high for SSE workload; each session holds memory for its duration Set MaxConcurrency = (instance_memory_MB − 300) / memory_per_session_MB
15–50s hang on first tool call after idle MinSize=0 — service scaled to zero, cold start on first request Set MinSize=1 for production, or use pause/resume schedule to pre-warm before business hours
CREATE_FAILED / UPDATE_FAILED on deployment Container fails health checks — application startup error, wrong port, missing env var Check /aws/apprunner/{name}/{id}/system log group for deployment event details
boto3 calls fail with NoCredentialsError inside container InstanceRoleArn not attached, or role trust policy uses wrong principal Trust principal must be tasks.apprunner.amazonaws.com; attach permissions in instance role (not access role)
Rollback to previous version impossible after bad deploy :latest ECR tag was used — previous image hash unknown Always tag production images with git SHA (main-{sha}); keep last 10 in ECR lifecycle policy
Rotated secret not picked up by running container App Runner containers do not hot-reload env vars; old containers keep cached secret value Call start_deployment() after rotating — forces new containers that load the updated value at startup
X-Ray traces show no downstream spans for boto3 calls patch_all() not called, or called after boto3 clients were already instantiated Call patch_all() at application start before any boto3 client creation
VPC connector update blocked — can't delete old connector version Service still references the old connector ARN Update service to reference new connector ARN first; wait for RUNNING; then delete old connector

Production checklists

Service creation checklist

VPC connector checklist

Observability checklist

CI/CD checklist

Monitor your App Runner MCP endpoints end-to-end

App Runner's internal observability covers what happens inside your service — but it misses the full network path from the internet. AliveMCP probes each MCP endpoint every 60 seconds from outside AWS, measuring end-to-end response time including DNS resolution, TLS handshake, and the complete path through App Runner's load balancer. CloudWatch metrics stay quiet during DNS propagation failures and App Runner load balancer degradation — external probing catches those before your users do.

Join the waitlist →