Guide · AWS CloudWatch · Alerting
CloudWatch Alarms for MCP Server Health Monitoring
CloudWatch metric alarms let you fire SNS notifications — and from there, Slack messages, webhook triggers, or PagerDuty pages — the moment your MCP server's error rate, throttle count, or p99 latency crosses a threshold you define. Composite alarms let you combine multiple signals (error rate AND high latency AND no throttles) into a single alert so you are not woken up by false positives. For querying historical log data after an incident, see CloudWatch Logs Insights for MCP server debugging; for structured application metrics emitted by your Lambda handlers, see CloudWatch Embedded Metric Format for MCP telemetry.
TL;DR
Create a CloudWatch alarm on the Lambda Errors metric for your MCP function, set the threshold to "Sum of Errors > 5 in 1 minute for 2 consecutive datapoints", and route it to an SNS topic that forwards to your Slack channel or webhook. Add a second alarm on Throttles and a third on the Duration p99 statistic. Combine all three into a composite alarm so a single ALARM state fires only when multiple signals are degraded simultaneously. Use INSUFFICIENT_DATA actions to detect when the function stops receiving traffic entirely.
The four CloudWatch metrics that matter most for MCP servers
Lambda publishes a standard set of metrics to CloudWatch under the AWS/Lambda namespace. For an MCP server, four metrics directly map to the health signals your users experience:
| Metric | What it measures | Alarm threshold recommendation |
|---|---|---|
Errors | Number of invocations that threw an unhandled exception or returned an error response | Sum > 5 in 1 minute; evaluate 2 consecutive datapoints before firing |
Throttles | Number of invocations rejected because concurrent execution limit was reached | Sum > 1 in 1 minute; throttles from a user's perspective are instant failures, so threshold is lower |
Duration | Wall-clock time for each invocation, in milliseconds | p99 > 5000ms (or 80% of your function timeout, whichever is lower) |
ConcurrentExecutions | Current number of simultaneous Lambda invocations | Max > 80% of your reserved concurrency limit |
If you have published custom metrics via CloudWatch Embedded Metric Format (EMF) or PutMetricData, you can alarm on those too — for example, an alarm on a custom McpToolErrors metric with dimensions by tool name, so you know which specific tool is failing rather than just that the overall function has errors.
Creating alarms via AWS CLI
# Alarm 1: MCP server error rate
# Fires if Lambda Errors sum exceeds 5 in a 1-minute window, for 2 consecutive windows
aws cloudwatch put-metric-alarm \
--alarm-name "mcp-server-prod-errors" \
--alarm-description "MCP server Lambda error rate is elevated" \
--namespace "AWS/Lambda" \
--metric-name "Errors" \
--dimensions Name=FunctionName,Value=mcp-server-prod \
--statistic Sum \
--period 60 \
--evaluation-periods 2 \
--threshold 5 \
--comparison-operator GreaterThanOrEqualToThreshold \
--treat-missing-data notBreaching \
--alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts \
--ok-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts
# Alarm 2: MCP server throttles
# Any throttle in a 1-minute window fires immediately (1 consecutive period)
aws cloudwatch put-metric-alarm \
--alarm-name "mcp-server-prod-throttles" \
--alarm-description "MCP server Lambda is throttling requests" \
--namespace "AWS/Lambda" \
--metric-name "Throttles" \
--dimensions Name=FunctionName,Value=mcp-server-prod \
--statistic Sum \
--period 60 \
--evaluation-periods 1 \
--threshold 1 \
--comparison-operator GreaterThanOrEqualToThreshold \
--treat-missing-data notBreaching \
--alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts
# Alarm 3: MCP server p99 latency
# p99 Duration exceeds 5 seconds for 3 consecutive minutes
aws cloudwatch put-metric-alarm \
--alarm-name "mcp-server-prod-latency-p99" \
--alarm-description "MCP server p99 latency is above 5s threshold" \
--namespace "AWS/Lambda" \
--metric-name "Duration" \
--dimensions Name=FunctionName,Value=mcp-server-prod \
--extended-statistic p99 \
--period 60 \
--evaluation-periods 3 \
--threshold 5000 \
--comparison-operator GreaterThanOrEqualToThreshold \
--treat-missing-data notBreaching \
--alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts
Composite alarms — reduce alert noise for MCP servers
A composite alarm evaluates a Boolean expression over the states of other alarms rather than over raw metric values. This is the correct tool for an MCP server because many degraded conditions are not individually alarming — a short error spike (errors alarm = ALARM) during a deploy is expected and not a real incident. But errors + high latency + no throttles simultaneously is a real incident.
# Composite alarm: fires only when BOTH error rate AND latency are degraded
# Uses alarm state expressions — references the other alarms by name
aws cloudwatch put-composite-alarm \
--alarm-name "mcp-server-prod-health" \
--alarm-description "MCP server is degraded: elevated errors AND high latency" \
--alarm-rule "ALARM(mcp-server-prod-errors) AND ALARM(mcp-server-prod-latency-p99)" \
--alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts-critical \
--ok-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts-critical
Composite alarm expressions support AND, OR, and NOT operators and can reference up to 100 component alarms. Useful patterns for MCP server observability:
# Degraded but not throttled (suspected bug, not capacity)
ALARM(mcp-server-prod-errors) AND ALARM(mcp-server-prod-latency-p99) AND NOT ALARM(mcp-server-prod-throttles)
# Throttled and not just a single spike (sustained capacity issue)
ALARM(mcp-server-prod-throttles) AND ALARM(mcp-server-prod-concurrency)
# Any signal is bad (broad alerting for low-traffic dev environments)
ALARM(mcp-server-dev-errors) OR ALARM(mcp-server-dev-throttles)
Anomaly detection alarms
For MCP servers with variable traffic patterns — higher invocation rates during business hours, lower at night — static thresholds on Errors or Duration can produce false positives. A 100 errors/minute alarm makes no sense for a server that normally handles 5,000 invocations/minute at 9 AM but only 50/minute at 2 AM. CloudWatch anomaly detection trains a statistical model on 2 weeks of metric history and expresses thresholds as bands around the expected value rather than fixed numbers.
# Create an anomaly detector for Duration p99
aws cloudwatch put-anomaly-detector \
--namespace "AWS/Lambda" \
--metric-name "Duration" \
--dimensions Name=FunctionName,Value=mcp-server-prod \
--stat "p99" \
--configuration ExcludedTimeRanges=[],MetricTimezone="UTC"
# Create an alarm that fires when Duration p99 is above the anomaly band (upper bound)
# The threshold is the width of the band in standard deviations; 2 = ~95% confidence interval
aws cloudwatch put-metric-alarm \
--alarm-name "mcp-server-prod-latency-anomaly" \
--alarm-description "MCP server p99 latency is anomalously high" \
--metrics '[
{
"Id": "m1",
"MetricStat": {
"Metric": {
"Namespace": "AWS/Lambda",
"MetricName": "Duration",
"Dimensions": [{"Name": "FunctionName", "Value": "mcp-server-prod"}]
},
"Period": 60,
"Stat": "p99"
}
},
{
"Id": "ad1",
"Expression": "ANOMALY_DETECTION_BAND(m1, 2)"
}
]' \
--comparison-operator GreaterThanUpperThreshold \
--threshold-metric-id ad1 \
--evaluation-periods 3 \
--treat-missing-data notBreaching \
--alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts
SNS → Slack webhook integration
CloudWatch alarms notify SNS topics, which can route to HTTP endpoints — including Slack incoming webhooks and any MCP server alert webhook you have configured. The simplest path is SNS → Lambda → Slack webhook (the Lambda function formats the message and calls Slack's API). For teams already using PagerDuty or OpsGenie, those services have native SNS integrations that auto-create incidents from alarm state changes.
# Create SNS topic for MCP alerts
aws sns create-topic --name mcp-alerts
# Subscribe a webhook Lambda to the topic
# (This Lambda receives the SNS notification and posts to Slack)
aws sns subscribe \
--topic-arn arn:aws:sns:us-east-1:123456789012:mcp-alerts \
--protocol lambda \
--notification-endpoint arn:aws:lambda:us-east-1:123456789012:function:mcp-slack-notifier
# The Lambda function receives events in this shape:
# {
# "Records": [{
# "Sns": {
# "Subject": "ALARM: \"mcp-server-prod-errors\" in US East (N. Virginia)",
# "Message": "{\"AlarmName\":\"mcp-server-prod-errors\",...}",
# ...
# }
# }]
# }
# Minimal Slack notifier Lambda (Node.js)
export const handler = async (event) => {
const snsMessage = JSON.parse(event.Records[0].Sns.Message);
const color = snsMessage.NewStateValue === 'ALARM' ? 'danger' : 'good';
await fetch(process.env.SLACK_WEBHOOK_URL, {
method: 'POST',
body: JSON.stringify({
attachments: [{
color,
title: snsMessage.AlarmName,
text: snsMessage.NewStateReason,
fields: [
{ title: 'State', value: snsMessage.NewStateValue, short: true },
{ title: 'Region', value: snsMessage.Region, short: true },
]
}]
})
});
};
CDK example — complete alarm stack for MCP servers
import * as cdk from 'aws-cdk-lib';
import * as cloudwatch from 'aws-cdk-lib/aws-cloudwatch';
import * as cloudwatch_actions from 'aws-cdk-lib/aws-cloudwatch-actions';
import * as sns from 'aws-cdk-lib/aws-sns';
import * as lambda from 'aws-cdk-lib/aws-lambda';
import { Construct } from 'constructs';
interface McpAlarmProps {
mcpFunction: lambda.IFunction;
alertTopic: sns.ITopic;
}
export class McpAlarmStack extends Construct {
constructor(scope: Construct, id: string, props: McpAlarmProps) {
super(scope, id);
const snsAction = new cloudwatch_actions.SnsAction(props.alertTopic);
// Error rate alarm — fires after 2 consecutive 1-min windows with 5+ errors
const errorsAlarm = new cloudwatch.Alarm(this, 'ErrorsAlarm', {
alarmName: `${props.mcpFunction.functionName}-errors`,
alarmDescription: 'MCP server error rate elevated',
metric: props.mcpFunction.metricErrors({
statistic: 'Sum',
period: cdk.Duration.minutes(1),
}),
threshold: 5,
evaluationPeriods: 2,
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
});
errorsAlarm.addAlarmAction(snsAction);
errorsAlarm.addOkAction(snsAction);
// Throttle alarm — fires immediately on any throttle
const throttlesAlarm = new cloudwatch.Alarm(this, 'ThrottlesAlarm', {
alarmName: `${props.mcpFunction.functionName}-throttles`,
alarmDescription: 'MCP server is throttling',
metric: props.mcpFunction.metricThrottles({
statistic: 'Sum',
period: cdk.Duration.minutes(1),
}),
threshold: 1,
evaluationPeriods: 1,
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
});
throttlesAlarm.addAlarmAction(snsAction);
// p99 latency alarm — fires after 3 consecutive minutes above 5s
const latencyAlarm = new cloudwatch.Alarm(this, 'LatencyAlarm', {
alarmName: `${props.mcpFunction.functionName}-latency-p99`,
alarmDescription: 'MCP server p99 latency above 5s',
metric: props.mcpFunction.metricDuration({
statistic: 'p99',
period: cdk.Duration.minutes(1),
}),
threshold: 5000,
evaluationPeriods: 3,
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
});
latencyAlarm.addAlarmAction(snsAction);
// Composite alarm — fires only when errors AND latency are both bad
const healthAlarm = new cloudwatch.CompositeAlarm(this, 'HealthAlarm', {
alarmRule: cloudwatch.AlarmRule.allOf(errorsAlarm, latencyAlarm),
compositeAlarmName: `${props.mcpFunction.functionName}-health`,
alarmDescription: 'MCP server health degraded: errors AND latency elevated',
});
healthAlarm.addAlarmAction(snsAction);
healthAlarm.addOkAction(snsAction);
}
}
AliveMCP + CloudWatch Alarms: external vs internal alerting
CloudWatch Alarms fire from metrics produced inside the AWS account — they are blind to failures that happen before a Lambda invocation is recorded, such as DNS resolution failures, TLS handshake errors, or API Gateway routing failures that reject requests before they reach your function. AliveMCP probes from outside your AWS account every 60 seconds, testing the full path from public internet to your MCP endpoint. Use AliveMCP for external reachability assurance and CloudWatch Alarms for internal health trending — both are needed for complete MCP server observability.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Alarm stays in INSUFFICIENT_DATA after creation | No data has been published to the metric yet (no Lambda invocations in the evaluation period), or the function name dimension is misspelled | Invoke the function once to generate a metric datapoint; verify the dimension value matches exactly: aws lambda get-function --function-name mcp-server-prod --query 'Configuration.FunctionName' |
| Composite alarm never fires even when component alarms are ALARM | The AlarmRule expression uses AND but one of the component alarms is INSUFFICIENT_DATA rather than ALARM; INSUFFICIENT_DATA is treated as neither OK nor ALARM in AND expressions | Add treat-missing-data missing to component alarms so they go to ALARM (not INSUFFICIENT_DATA) when data stops flowing |
| p99 statistic not available in alarm creation | Standard statistics (--statistic) only support Average, Sum, SampleCount, Min, Max; percentiles require the --extended-statistic p99 parameter | Use --extended-statistic p99 instead of --statistic p99 in the CLI; in CDK use statistic: 'p99' string literal (not the cloudwatch.Stats.percentile(99) method, which produces the wrong format for Lambda metrics) |
| Alarm firing constantly during deploys | A Lambda function update causes a brief period of elevated errors (new code path hits uncaught exceptions) that breaches the threshold on every deploy | Add a suppression window: use composite alarm with a "deploy in progress" alarm (triggered by a CodeDeploy deployment metric) to mask the component alarms during the deployment window |
| SNS notification not received | SNS topic policy does not allow CloudWatch to publish; or the Lambda subscriber has no permission to be invoked by SNS | Check SNS topic policy for sns:Publish from cloudwatch.amazonaws.com; check Lambda resource policy for lambda:InvokeFunction from sns.amazonaws.com |