Guide · AWS CloudWatch · Alerting

CloudWatch Alarms for MCP Server Health Monitoring

CloudWatch metric alarms let you fire SNS notifications — and from there, Slack messages, webhook triggers, or PagerDuty pages — the moment your MCP server's error rate, throttle count, or p99 latency crosses a threshold you define. Composite alarms let you combine multiple signals (error rate AND high latency AND no throttles) into a single alert so you are not woken up by false positives. For querying historical log data after an incident, see CloudWatch Logs Insights for MCP server debugging; for structured application metrics emitted by your Lambda handlers, see CloudWatch Embedded Metric Format for MCP telemetry.

TL;DR

Create a CloudWatch alarm on the Lambda Errors metric for your MCP function, set the threshold to "Sum of Errors > 5 in 1 minute for 2 consecutive datapoints", and route it to an SNS topic that forwards to your Slack channel or webhook. Add a second alarm on Throttles and a third on the Duration p99 statistic. Combine all three into a composite alarm so a single ALARM state fires only when multiple signals are degraded simultaneously. Use INSUFFICIENT_DATA actions to detect when the function stops receiving traffic entirely.

The four CloudWatch metrics that matter most for MCP servers

Lambda publishes a standard set of metrics to CloudWatch under the AWS/Lambda namespace. For an MCP server, four metrics directly map to the health signals your users experience:

MetricWhat it measuresAlarm threshold recommendation
ErrorsNumber of invocations that threw an unhandled exception or returned an error responseSum > 5 in 1 minute; evaluate 2 consecutive datapoints before firing
ThrottlesNumber of invocations rejected because concurrent execution limit was reachedSum > 1 in 1 minute; throttles from a user's perspective are instant failures, so threshold is lower
DurationWall-clock time for each invocation, in millisecondsp99 > 5000ms (or 80% of your function timeout, whichever is lower)
ConcurrentExecutionsCurrent number of simultaneous Lambda invocationsMax > 80% of your reserved concurrency limit

If you have published custom metrics via CloudWatch Embedded Metric Format (EMF) or PutMetricData, you can alarm on those too — for example, an alarm on a custom McpToolErrors metric with dimensions by tool name, so you know which specific tool is failing rather than just that the overall function has errors.

Creating alarms via AWS CLI

# Alarm 1: MCP server error rate
# Fires if Lambda Errors sum exceeds 5 in a 1-minute window, for 2 consecutive windows
aws cloudwatch put-metric-alarm \
  --alarm-name "mcp-server-prod-errors" \
  --alarm-description "MCP server Lambda error rate is elevated" \
  --namespace "AWS/Lambda" \
  --metric-name "Errors" \
  --dimensions Name=FunctionName,Value=mcp-server-prod \
  --statistic Sum \
  --period 60 \
  --evaluation-periods 2 \
  --threshold 5 \
  --comparison-operator GreaterThanOrEqualToThreshold \
  --treat-missing-data notBreaching \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts \
  --ok-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts

# Alarm 2: MCP server throttles
# Any throttle in a 1-minute window fires immediately (1 consecutive period)
aws cloudwatch put-metric-alarm \
  --alarm-name "mcp-server-prod-throttles" \
  --alarm-description "MCP server Lambda is throttling requests" \
  --namespace "AWS/Lambda" \
  --metric-name "Throttles" \
  --dimensions Name=FunctionName,Value=mcp-server-prod \
  --statistic Sum \
  --period 60 \
  --evaluation-periods 1 \
  --threshold 1 \
  --comparison-operator GreaterThanOrEqualToThreshold \
  --treat-missing-data notBreaching \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts

# Alarm 3: MCP server p99 latency
# p99 Duration exceeds 5 seconds for 3 consecutive minutes
aws cloudwatch put-metric-alarm \
  --alarm-name "mcp-server-prod-latency-p99" \
  --alarm-description "MCP server p99 latency is above 5s threshold" \
  --namespace "AWS/Lambda" \
  --metric-name "Duration" \
  --dimensions Name=FunctionName,Value=mcp-server-prod \
  --extended-statistic p99 \
  --period 60 \
  --evaluation-periods 3 \
  --threshold 5000 \
  --comparison-operator GreaterThanOrEqualToThreshold \
  --treat-missing-data notBreaching \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts

Composite alarms — reduce alert noise for MCP servers

A composite alarm evaluates a Boolean expression over the states of other alarms rather than over raw metric values. This is the correct tool for an MCP server because many degraded conditions are not individually alarming — a short error spike (errors alarm = ALARM) during a deploy is expected and not a real incident. But errors + high latency + no throttles simultaneously is a real incident.

# Composite alarm: fires only when BOTH error rate AND latency are degraded
# Uses alarm state expressions — references the other alarms by name
aws cloudwatch put-composite-alarm \
  --alarm-name "mcp-server-prod-health" \
  --alarm-description "MCP server is degraded: elevated errors AND high latency" \
  --alarm-rule "ALARM(mcp-server-prod-errors) AND ALARM(mcp-server-prod-latency-p99)" \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts-critical \
  --ok-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts-critical

Composite alarm expressions support AND, OR, and NOT operators and can reference up to 100 component alarms. Useful patterns for MCP server observability:

# Degraded but not throttled (suspected bug, not capacity)
ALARM(mcp-server-prod-errors) AND ALARM(mcp-server-prod-latency-p99) AND NOT ALARM(mcp-server-prod-throttles)

# Throttled and not just a single spike (sustained capacity issue)
ALARM(mcp-server-prod-throttles) AND ALARM(mcp-server-prod-concurrency)

# Any signal is bad (broad alerting for low-traffic dev environments)
ALARM(mcp-server-dev-errors) OR ALARM(mcp-server-dev-throttles)

Anomaly detection alarms

For MCP servers with variable traffic patterns — higher invocation rates during business hours, lower at night — static thresholds on Errors or Duration can produce false positives. A 100 errors/minute alarm makes no sense for a server that normally handles 5,000 invocations/minute at 9 AM but only 50/minute at 2 AM. CloudWatch anomaly detection trains a statistical model on 2 weeks of metric history and expresses thresholds as bands around the expected value rather than fixed numbers.

# Create an anomaly detector for Duration p99
aws cloudwatch put-anomaly-detector \
  --namespace "AWS/Lambda" \
  --metric-name "Duration" \
  --dimensions Name=FunctionName,Value=mcp-server-prod \
  --stat "p99" \
  --configuration ExcludedTimeRanges=[],MetricTimezone="UTC"

# Create an alarm that fires when Duration p99 is above the anomaly band (upper bound)
# The threshold is the width of the band in standard deviations; 2 = ~95% confidence interval
aws cloudwatch put-metric-alarm \
  --alarm-name "mcp-server-prod-latency-anomaly" \
  --alarm-description "MCP server p99 latency is anomalously high" \
  --metrics '[
    {
      "Id": "m1",
      "MetricStat": {
        "Metric": {
          "Namespace": "AWS/Lambda",
          "MetricName": "Duration",
          "Dimensions": [{"Name": "FunctionName", "Value": "mcp-server-prod"}]
        },
        "Period": 60,
        "Stat": "p99"
      }
    },
    {
      "Id": "ad1",
      "Expression": "ANOMALY_DETECTION_BAND(m1, 2)"
    }
  ]' \
  --comparison-operator GreaterThanUpperThreshold \
  --threshold-metric-id ad1 \
  --evaluation-periods 3 \
  --treat-missing-data notBreaching \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts

SNS → Slack webhook integration

CloudWatch alarms notify SNS topics, which can route to HTTP endpoints — including Slack incoming webhooks and any MCP server alert webhook you have configured. The simplest path is SNS → Lambda → Slack webhook (the Lambda function formats the message and calls Slack's API). For teams already using PagerDuty or OpsGenie, those services have native SNS integrations that auto-create incidents from alarm state changes.

# Create SNS topic for MCP alerts
aws sns create-topic --name mcp-alerts

# Subscribe a webhook Lambda to the topic
# (This Lambda receives the SNS notification and posts to Slack)
aws sns subscribe \
  --topic-arn arn:aws:sns:us-east-1:123456789012:mcp-alerts \
  --protocol lambda \
  --notification-endpoint arn:aws:lambda:us-east-1:123456789012:function:mcp-slack-notifier

# The Lambda function receives events in this shape:
# {
#   "Records": [{
#     "Sns": {
#       "Subject": "ALARM: \"mcp-server-prod-errors\" in US East (N. Virginia)",
#       "Message": "{\"AlarmName\":\"mcp-server-prod-errors\",...}",
#       ...
#     }
#   }]
# }

# Minimal Slack notifier Lambda (Node.js)
export const handler = async (event) => {
  const snsMessage = JSON.parse(event.Records[0].Sns.Message);
  const color = snsMessage.NewStateValue === 'ALARM' ? 'danger' : 'good';
  await fetch(process.env.SLACK_WEBHOOK_URL, {
    method: 'POST',
    body: JSON.stringify({
      attachments: [{
        color,
        title: snsMessage.AlarmName,
        text: snsMessage.NewStateReason,
        fields: [
          { title: 'State', value: snsMessage.NewStateValue, short: true },
          { title: 'Region', value: snsMessage.Region, short: true },
        ]
      }]
    })
  });
};

CDK example — complete alarm stack for MCP servers

import * as cdk from 'aws-cdk-lib';
import * as cloudwatch from 'aws-cdk-lib/aws-cloudwatch';
import * as cloudwatch_actions from 'aws-cdk-lib/aws-cloudwatch-actions';
import * as sns from 'aws-cdk-lib/aws-sns';
import * as lambda from 'aws-cdk-lib/aws-lambda';
import { Construct } from 'constructs';

interface McpAlarmProps {
  mcpFunction: lambda.IFunction;
  alertTopic: sns.ITopic;
}

export class McpAlarmStack extends Construct {
  constructor(scope: Construct, id: string, props: McpAlarmProps) {
    super(scope, id);

    const snsAction = new cloudwatch_actions.SnsAction(props.alertTopic);

    // Error rate alarm — fires after 2 consecutive 1-min windows with 5+ errors
    const errorsAlarm = new cloudwatch.Alarm(this, 'ErrorsAlarm', {
      alarmName: `${props.mcpFunction.functionName}-errors`,
      alarmDescription: 'MCP server error rate elevated',
      metric: props.mcpFunction.metricErrors({
        statistic: 'Sum',
        period: cdk.Duration.minutes(1),
      }),
      threshold: 5,
      evaluationPeriods: 2,
      treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
    });
    errorsAlarm.addAlarmAction(snsAction);
    errorsAlarm.addOkAction(snsAction);

    // Throttle alarm — fires immediately on any throttle
    const throttlesAlarm = new cloudwatch.Alarm(this, 'ThrottlesAlarm', {
      alarmName: `${props.mcpFunction.functionName}-throttles`,
      alarmDescription: 'MCP server is throttling',
      metric: props.mcpFunction.metricThrottles({
        statistic: 'Sum',
        period: cdk.Duration.minutes(1),
      }),
      threshold: 1,
      evaluationPeriods: 1,
      treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
    });
    throttlesAlarm.addAlarmAction(snsAction);

    // p99 latency alarm — fires after 3 consecutive minutes above 5s
    const latencyAlarm = new cloudwatch.Alarm(this, 'LatencyAlarm', {
      alarmName: `${props.mcpFunction.functionName}-latency-p99`,
      alarmDescription: 'MCP server p99 latency above 5s',
      metric: props.mcpFunction.metricDuration({
        statistic: 'p99',
        period: cdk.Duration.minutes(1),
      }),
      threshold: 5000,
      evaluationPeriods: 3,
      treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
    });
    latencyAlarm.addAlarmAction(snsAction);

    // Composite alarm — fires only when errors AND latency are both bad
    const healthAlarm = new cloudwatch.CompositeAlarm(this, 'HealthAlarm', {
      alarmRule: cloudwatch.AlarmRule.allOf(errorsAlarm, latencyAlarm),
      compositeAlarmName: `${props.mcpFunction.functionName}-health`,
      alarmDescription: 'MCP server health degraded: errors AND latency elevated',
    });
    healthAlarm.addAlarmAction(snsAction);
    healthAlarm.addOkAction(snsAction);
  }
}

AliveMCP + CloudWatch Alarms: external vs internal alerting

CloudWatch Alarms fire from metrics produced inside the AWS account — they are blind to failures that happen before a Lambda invocation is recorded, such as DNS resolution failures, TLS handshake errors, or API Gateway routing failures that reject requests before they reach your function. AliveMCP probes from outside your AWS account every 60 seconds, testing the full path from public internet to your MCP endpoint. Use AliveMCP for external reachability assurance and CloudWatch Alarms for internal health trending — both are needed for complete MCP server observability.

Failure modes

SymptomCauseFix
Alarm stays in INSUFFICIENT_DATA after creationNo data has been published to the metric yet (no Lambda invocations in the evaluation period), or the function name dimension is misspelledInvoke the function once to generate a metric datapoint; verify the dimension value matches exactly: aws lambda get-function --function-name mcp-server-prod --query 'Configuration.FunctionName'
Composite alarm never fires even when component alarms are ALARMThe AlarmRule expression uses AND but one of the component alarms is INSUFFICIENT_DATA rather than ALARM; INSUFFICIENT_DATA is treated as neither OK nor ALARM in AND expressionsAdd treat-missing-data missing to component alarms so they go to ALARM (not INSUFFICIENT_DATA) when data stops flowing
p99 statistic not available in alarm creationStandard statistics (--statistic) only support Average, Sum, SampleCount, Min, Max; percentiles require the --extended-statistic p99 parameterUse --extended-statistic p99 instead of --statistic p99 in the CLI; in CDK use statistic: 'p99' string literal (not the cloudwatch.Stats.percentile(99) method, which produces the wrong format for Lambda metrics)
Alarm firing constantly during deploysA Lambda function update causes a brief period of elevated errors (new code path hits uncaught exceptions) that breaches the threshold on every deployAdd a suppression window: use composite alarm with a "deploy in progress" alarm (triggered by a CodeDeploy deployment metric) to mask the component alarms during the deployment window
SNS notification not receivedSNS topic policy does not allow CloudWatch to publish; or the Lambda subscriber has no permission to be invoked by SNSCheck SNS topic policy for sns:Publish from cloudwatch.amazonaws.com; check Lambda resource policy for lambda:InvokeFunction from sns.amazonaws.com