Guide · AWS App Runner · Auto Scaling · SSE Streaming · MCP Servers

App Runner Auto Scaling for MCP Servers with SSE Streaming

App Runner scales by measuring concurrent requests per instance — and for MCP servers using Server-Sent Events (SSE), one active MCP session holds one open HTTP connection for its entire duration, counting as one concurrent request the entire time. This makes MaxConcurrency the single most important auto-scaling parameter for SSE-based MCP servers: if MaxConcurrency is set too high (e.g., the default of 100), App Runner may run dozens of SSE sessions on a single instance before adding another, causing memory and CPU saturation; if set too low, App Runner adds instances aggressively for every few sessions, driving up compute cost. The right value depends on how much memory and CPU each MCP session consumes. Three decisions drive the auto-scaling configuration: MaxConcurrency for SSE (typically 10–50 for memory-intensive tools, 50–100 for lightweight ones), MinSize=0 vs MinSize=1 (scale-to-zero saves cost but adds 10–30s cold starts), and pause/resume scheduling for predictable off-hours traffic patterns.

TL;DR

Create an AutoScalingConfiguration with MaxConcurrency=25 for SSE-heavy MCP servers (each SSE session counts as 1 concurrent request for its full duration). Set MinSize=1 to eliminate cold starts for production services. Use MaxSize to cap cost. For dev/staging environments, set MinSize=0 and accept the cold start. Pause the service during off-hours with pause_service() to stop all instance charges.

Creating a custom auto-scaling configuration

App Runner's default auto-scaling configuration uses MaxConcurrency=100 and MinSize=1 / MaxSize=25. For MCP servers, override these with a custom configuration:

import boto3

apprunner = boto3.client("apprunner", region_name="us-east-1")

# Create a named auto-scaling config for SSE-based MCP servers
asc_response = apprunner.create_auto_scaling_configuration(
    AutoScalingConfigurationName="mcp-server-sse-scaling",
    MaxConcurrency=25,   # add a new instance when any existing instance hits 25 concurrent connections
    MinSize=1,           # always keep at least 1 instance running (no cold starts)
    MaxSize=10,          # never exceed 10 instances (cost cap)
)
asc_arn = asc_response["AutoScalingConfiguration"]["AutoScalingConfigurationArn"]

# Create the App Runner service using the custom scaling config
apprunner.create_service(
    ServiceName="mcp-server-prod",
    AutoScalingConfigurationArn=asc_arn,
    # ... (SourceConfiguration, InstanceConfiguration, NetworkConfiguration)
)

# Or attach to an existing service
apprunner.update_service(
    ServiceArn="arn:aws:apprunner:us-east-1:123456789012:service/mcp-server-prod/...",
    AutoScalingConfigurationArn=asc_arn,
)

The auto-scaling configuration is versioned. Creating a new configuration with the same name increments the version — old versions persist until deleted and can be re-attached to services.

MaxConcurrency for SSE vs HTTP request/response MCP patterns

The right MaxConcurrency value differs significantly between SSE and request/response MCP transports:

# SSE transport (HTTP streaming):
# - Each MCP session = 1 open HTTP connection = 1 concurrent request
# - Duration: minutes to hours (entire MCP session lifespan)
# - Memory per session: depends on tool state (10–200 MB per session)
# - Typical MaxConcurrency: 10–50
#   - 10: memory-heavy sessions (embedding models, in-memory caches)
#   - 25: standard stateful tools (DB connections, API clients)
#   - 50: lightweight stateless tools (quick API calls)

# HTTP + JSON-RPC transport (request/response):
# - Each MCP call = 1 request with short duration (50ms–5s)
# - Connection released immediately after response
# - MaxConcurrency measures peak in-flight calls, not session count
# - Typical MaxConcurrency: 50–100 (standard HTTP API behavior)

# Calculating MaxConcurrency for your SSE workload:
#   instance_memory_gb = 2.0   # from InstanceConfiguration
#   reserved_for_runtime_gb = 0.3  # OS + App Runner agent + Python/Node overhead
#   available_for_sessions_gb = instance_memory_gb - reserved_for_runtime_gb  # 1.7 GB
#   memory_per_session_mb = 80  # measure via CloudWatch MemoryUtilization
#   max_sessions_per_instance = (available_for_sessions_gb * 1024) / memory_per_session_mb
#   # = 1700 / 80 = ~21 sessions per instance
#   # Set MaxConcurrency = 20 (slight buffer below the calculated limit)

# Monitor actual memory usage to calibrate:
# CloudWatch metric: AppRunner/MemoryUtilization for the service
# If utilization consistently > 80% per instance, lower MaxConcurrency
# If utilization consistently < 40% per instance, raise MaxConcurrency

Scale-to-zero: MinSize=0 cold start behavior

Setting MinSize=0 allows App Runner to shut down all instances when there are no active requests, eliminating instance charges during idle periods. The downside is a cold start when the first request arrives after the service has scaled to zero:

# Cold start timeline (MinSize=0):
# 1. Request arrives (first MCP connection after idle period)
# 2. App Runner detects zero running instances, starts provisioning
# 3. ECR image pull: 5–15s (depends on image size; cached layers help)
# 4. Container start: 5–15s (JVM/Python interpreter init + app startup)
# 5. Health check passes: 5–20s (depends on health check Interval/Threshold)
# 6. Request routed to the new instance: total 15–50s cold start
#
# For MCP clients: the first tool call after a cold start will hang
# for 15–50 seconds before responding — often mistaken for a network error

# Minimize cold start time:
# 1. Keep the Docker image small (use slim base images, multi-stage builds)
# 2. Minimize startup work (lazy-init DB pools, defer heavy imports)
# 3. Add a readiness delay to the health check response (503 until ready)
#    so App Runner doesn't route before the app is actually ready

# Python: reduce import time via lazy loading
def get_heavy_client():
    # Don't import at module level — delay until first use
    import anthropic  # only imported when first tool call arrives
    return anthropic.Anthropic()

# Node.js: defer connection pool creation
let dbPool = null;
function getPool() {
    if (!dbPool) dbPool = createPool(process.env.DATABASE_URL);
    return dbPool;
}

Use MinSize=0 for: development and staging environments, internal tools with predictable daily patterns (business hours only), and demo deployments where cost matters more than latency. Use MinSize=1 for production MCP servers where MCP clients expect sub-second response times.

Pause and resume for scheduled off-hours cost reduction

Pausing an App Runner service stops all running instances and halts instance charges. The service configuration is preserved — resuming re-deploys the same image immediately:

import boto3
from datetime import datetime, timezone

apprunner = boto3.client("apprunner", region_name="us-east-1")

SERVICE_ARN = "arn:aws:apprunner:us-east-1:123456789012:service/mcp-server-dev/abc123"

def pause_service():
    """Stop all instances — used for off-hours cost savings."""
    apprunner.pause_service(ServiceArn=SERVICE_ARN)
    print(f"Service paused at {datetime.now(timezone.utc).isoformat()}")

def resume_service():
    """Restart instances — call before business hours begin."""
    apprunner.resume_service(ServiceArn=SERVICE_ARN)
    print(f"Service resuming at {datetime.now(timezone.utc).isoformat()}")

# Service states:
# RUNNING → pause_service() → OPERATION_IN_PROGRESS → PAUSED
# PAUSED → resume_service() → OPERATION_IN_PROGRESS → RUNNING
#
# During PAUSED state:
# - No instance charges (only service configuration is retained)
# - Requests return 503 Service Unavailable
# - Service URL remains registered (no DNS changes)
# - Resume time: 30–90s (image pull + container start + health check)

# Automate with AWS EventBridge Scheduler (cron expression):
# Pause: cron(0 22 * * ? *)  → 10 PM UTC daily
# Resume: cron(0 8 * * ? *)  → 8 AM UTC daily

For production MCP servers with business-hours traffic patterns, pause/resume can reduce monthly compute costs by 60–70% (e.g., 16 off-hours per day). The pause/resume approach is more cost-effective than scaling to zero with MinSize=0 for predictable patterns, because it avoids cold start latency during business hours when the service resumes at a known time.

Observing auto-scaling activity via CloudWatch metrics

App Runner emits auto-scaling metrics to CloudWatch that show when instances were added or removed:

import boto3

cloudwatch = boto3.client("cloudwatch", region_name="us-east-1")

# Key metrics for auto-scaling observation
metrics = [
    {
        "name": "ActiveInstances",
        "stat": "Maximum",
        "description": "Number of instances App Runner is billing (includes startup/shutdown)"
    },
    {
        "name": "Concurrency",
        "stat": "Maximum",
        "description": "Peak concurrent requests across all instances in the period"
    },
    {
        "name": "RequestLatency",
        "stat": "p99",
        "description": "P99 request duration — spikes during scale-out (new instance warming)"
    },
]

for metric in metrics:
    response = cloudwatch.get_metric_statistics(
        Namespace="AWS/AppRunner",
        MetricName=metric["name"],
        Dimensions=[{"Name": "ServiceName", "Value": "mcp-server-prod"}],
        StartTime="2026-10-01T00:00:00Z",
        EndTime="2026-10-01T23:59:59Z",
        Period=300,  # 5-minute granularity
        Statistics=[metric["stat"]],
    )
    datapoints = sorted(response["Datapoints"], key=lambda d: d["Timestamp"])
    for dp in datapoints[-5:]:  # last 5 datapoints
        print(f"{dp['Timestamp']}: {metric['name']}={dp[metric['stat']]}")

Watch for Concurrency approaching MaxConcurrency × ActiveInstances — this indicates the service is near its scaling threshold. Sustained P99 latency spikes correlated with instance count increases indicate that scale-out is happening too late; lower MaxConcurrency to trigger earlier scale-out.

Know when your App Runner MCP service scales down unexpectedly

App Runner scale-to-zero with MinSize=0 means MCP clients can hit cold starts without any warning. AliveMCP monitors MCP endpoint availability and response time every 60 seconds — alerting you when cold starts or scale-down events are degrading the tool-calling experience for users.

Join the waitlist →