Guide · AWS App Runner · CloudWatch · X-Ray · MCP Server Observability

App Runner MCP Server Observability — CloudWatch Logs, Metrics, and X-Ray

App Runner automatically sends application logs to CloudWatch Logs and emits service metrics to CloudWatch Metrics — no agents to install, no log shippers to configure. For MCP servers, the key observability concern is linking MCP tool call failures (a tool returning an error) to the underlying infrastructure events (scale-out latency, database connection exhaustion, downstream API rate limits). App Runner's built-in observability covers the infrastructure layer; structured application logging and X-Ray tracing cover the tool execution layer. Three components complete the picture: structured log output (JSON logs with tool name, session ID, duration, and error fields that can be filtered in CloudWatch Logs Insights), alarms on 5xx rate and P99 latency (the two leading indicators of MCP tool degradation), and X-Ray tracing (traces the full call chain from MCP client through App Runner through downstream AWS service calls).

TL;DR

App Runner sends stdout/stderr to /aws/apprunner/{service_name}/{service_id}/application automatically. Key metrics to alarm on: 5xxStatusResponses (any value > 0 = broken tools) and RequestLatency p99 (spike = scale-out or downstream latency). Enable X-Ray by creating an ObservabilityConfiguration with TraceConfiguration.Vendor=AWSXRAY and attaching it to the service.

CloudWatch log groups for App Runner services

App Runner creates two log groups automatically when a service is created:

# Application log group — contains stdout/stderr from your container
# /aws/apprunner/{service_name}/{service_id}/application
#
# System log group — App Runner platform events (deployments, health checks, scaling)
# /aws/apprunner/{service_name}/{service_id}/system

import boto3

logs = boto3.client("logs", region_name="us-east-1")

SERVICE_NAME = "mcp-server-prod"
SERVICE_ID = "abc123def456"  # from DescribeService response

APP_LOG_GROUP = f"/aws/apprunner/{SERVICE_NAME}/{SERVICE_ID}/application"
SYS_LOG_GROUP = f"/aws/apprunner/{SERVICE_NAME}/{SERVICE_ID}/system"

# Query recent application logs
response = logs.filter_log_events(
    logGroupName=APP_LOG_GROUP,
    startTime=int((time.time() - 3600) * 1000),  # last 1 hour
    filterPattern='{ $.level = "ERROR" }',  # JSON structured log filter
    limit=50,
)

for event in response["events"]:
    import json
    try:
        entry = json.loads(event["message"])
        print(f"{entry['timestamp']} [{entry['tool']}] {entry['error']}")
    except json.JSONDecodeError:
        print(event["message"])

The system log group contains deployment events, health check failures, and scaling activity — invaluable for diagnosing why a new deployment rolled back or why instances are being replaced during an active traffic period.

Structured logging from MCP server tools

App Runner captures everything written to stdout/stderr. Using structured JSON logging makes logs queryable in CloudWatch Logs Insights without parsing regex:

import json
import time
import sys
import traceback
from contextvars import ContextVar

# Context variables injected at session start
session_id: ContextVar[str] = ContextVar("session_id", default="unknown")
client_id: ContextVar[str] = ContextVar("client_id", default="unknown")

def log(level: str, tool: str, message: str, **extra):
    """Emit a JSON log line to stdout — captured by App Runner → CloudWatch."""
    entry = {
        "timestamp": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "level": level,
        "tool": tool,
        "message": message,
        "session_id": session_id.get(),
        "client_id": client_id.get(),
        **extra,
    }
    print(json.dumps(entry), flush=True)  # flush=True for immediate CloudWatch delivery

# Decorator to log every MCP tool call
def logged_tool(tool_name: str):
    def decorator(fn):
        async def wrapper(*args, **kwargs):
            start = time.monotonic()
            log("INFO", tool_name, "tool_call_start", args_count=len(args))
            try:
                result = await fn(*args, **kwargs)
                duration_ms = int((time.monotonic() - start) * 1000)
                log("INFO", tool_name, "tool_call_success", duration_ms=duration_ms)
                return result
            except Exception as e:
                duration_ms = int((time.monotonic() - start) * 1000)
                log("ERROR", tool_name, "tool_call_error",
                    error=str(e),
                    error_type=type(e).__name__,
                    duration_ms=duration_ms,
                    traceback=traceback.format_exc())
                raise
        return wrapper
    return decorator

@logged_tool("fetch_workspace")
async def fetch_workspace(workspace_id: str) -> dict:
    # MCP tool implementation
    ...

CloudWatch Logs Insights queries for MCP tools

Logs Insights provides SQL-like queries over CloudWatch log groups. Useful queries for MCP server debugging:

# Error rate by tool — which tools fail most?
fields @timestamp, tool, error_type
| filter level = "ERROR"
| stats count(*) as error_count by tool, error_type
| sort error_count desc
| limit 20

# P95 latency by tool — which tools are slow?
fields @timestamp, tool, duration_ms
| filter level = "INFO" and message = "tool_call_success"
| stats
    pct(duration_ms, 50) as p50_ms,
    pct(duration_ms, 95) as p95_ms,
    pct(duration_ms, 99) as p99_ms,
    count(*) as call_count
  by tool
| sort p99_ms desc

# Sessions with errors in the last hour — incident investigation
fields @timestamp, session_id, tool, error, error_type
| filter level = "ERROR"
| sort @timestamp desc
| limit 100

# Tool call volume by hour — traffic patterns
fields @timestamp, tool
| filter message = "tool_call_start"
| stats count(*) as calls by bin(1h), tool
| sort @timestamp asc

Run these queries via CloudWatch console or boto3's start_query() API. Results are available within 1–2 seconds for log groups with less than 5 GB of data in the time range queried.

CloudWatch alarms for MCP service health

Set up alarms on the two most actionable App Runner metrics for MCP servers:

import boto3

cloudwatch = boto3.client("cloudwatch", region_name="us-east-1")

SERVICE_NAME = "mcp-server-prod"
SNS_TOPIC_ARN = "arn:aws:sns:us-east-1:123456789012:mcp-alerts"

# Alarm 1: Any 5xx responses — immediate action required
cloudwatch.put_metric_alarm(
    AlarmName=f"{SERVICE_NAME}-5xx-errors",
    AlarmDescription="App Runner returning 5xx — MCP tools broken",
    Namespace="AWS/AppRunner",
    MetricName="5xxStatusResponses",
    Dimensions=[{"Name": "ServiceName", "Value": SERVICE_NAME}],
    Statistic="Sum",
    Period=60,           # 1-minute evaluation
    EvaluationPeriods=1,
    Threshold=1,         # trigger on any 5xx
    ComparisonOperator="GreaterThanOrEqualToThreshold",
    TreatMissingData="notBreaching",
    AlarmActions=[SNS_TOPIC_ARN],
    OKActions=[SNS_TOPIC_ARN],
)

# Alarm 2: High P99 latency — MCP tool slowdown
cloudwatch.put_metric_alarm(
    AlarmName=f"{SERVICE_NAME}-high-latency",
    AlarmDescription="App Runner P99 latency > 5s — MCP tools degraded",
    Namespace="AWS/AppRunner",
    MetricName="RequestLatency",
    Dimensions=[{"Name": "ServiceName", "Value": SERVICE_NAME}],
    ExtendedStatistic="p99",
    Period=300,          # 5-minute evaluation
    EvaluationPeriods=2, # 2 consecutive periods = 10 minutes of high latency
    Threshold=5000,      # 5000ms = 5 seconds
    ComparisonOperator="GreaterThanThreshold",
    TreatMissingData="notBreaching",
    AlarmActions=[SNS_TOPIC_ARN],
    OKActions=[SNS_TOPIC_ARN],
)

# Alarm 3: Instance count spike — unexpected scaling (could mean traffic surge or flapping)
cloudwatch.put_metric_alarm(
    AlarmName=f"{SERVICE_NAME}-instance-spike",
    AlarmDescription="App Runner instance count > 5 — unexpected traffic surge",
    Namespace="AWS/AppRunner",
    MetricName="ActiveInstances",
    Dimensions=[{"Name": "ServiceName", "Value": SERVICE_NAME}],
    Statistic="Maximum",
    Period=300,
    EvaluationPeriods=1,
    Threshold=5,
    ComparisonOperator="GreaterThanThreshold",
    TreatMissingData="notBreaching",
    AlarmActions=[SNS_TOPIC_ARN],
)

Enabling X-Ray tracing for MCP tool call chains

X-Ray tracing captures the full span of a request through App Runner and into downstream AWS services (DynamoDB, S3, Bedrock). Enable it via an observability configuration:

import boto3

apprunner = boto3.client("apprunner", region_name="us-east-1")

# Step 1: Create an observability configuration with X-Ray enabled
obs_response = apprunner.create_observability_configuration(
    ObservabilityConfigurationName="mcp-server-xray",
    TraceConfiguration={
        "Vendor": "AWSXRAY",  # the only supported vendor
    },
)
obs_arn = obs_response["ObservabilityConfiguration"]["ObservabilityConfigurationArn"]

# Step 2: Attach to the service
apprunner.update_service(
    ServiceArn="arn:aws:apprunner:us-east-1:123456789012:service/mcp-server-prod/...",
    ObservabilityConfiguration={
        "ObservabilityEnabled": True,
        "ObservabilityConfigurationArn": obs_arn,
    },
)

# Step 3: In the container, use aws_xray_sdk to create subsegments per tool
from aws_xray_sdk.core import xray_recorder
from aws_xray_sdk.core import patch_all

patch_all()  # auto-instrument boto3, requests, aiohttp, etc.

async def my_mcp_tool(args: dict) -> str:
    with xray_recorder.in_subsegment("mcp_tool_execution") as subsegment:
        subsegment.put_annotation("tool_name", "fetch_workspace")
        subsegment.put_annotation("session_id", session_id.get())
        subsegment.put_metadata("args", args)

        # boto3 calls inside are auto-traced as child subsegments
        result = await fetch_from_dynamodb(args["workspace_id"])

        subsegment.put_metadata("result_size", len(str(result)))
        return result

App Runner injects the X-Ray daemon as a sidecar — the SDK sends UDP trace data to 127.0.0.1:2000 (the daemon's default address), and the daemon batches and forwards traces to the X-Ray service. The patch_all() call automatically wraps boto3 clients so every DynamoDB, S3, and Bedrock call appears as a child segment in the trace.

External monitoring for App Runner MCP endpoints

CloudWatch observability covers what happens inside your App Runner service. AliveMCP provides external-perspective monitoring — probing the MCP endpoint from outside AWS every 60 seconds, measuring end-to-end response time including DNS resolution, TLS handshake, and the full network path from the internet to your container. Catches issues that internal CloudWatch metrics miss, like DNS propagation failures or App Runner load balancer problems.

Join the waitlist →