Guide · EventBridge Scheduler · Observability

Monitoring EventBridge Scheduler for MCP Servers — CloudWatch Metrics, DLQ Alarms, Audit Trail

EventBridge Scheduler publishes six CloudWatch metrics under the AWS/Scheduler namespace, but two are critical for MCP server monitoring health: InvocationAttemptCount (total invocations attempted) and InvocationDroppedCount (invocations that exhausted retries and were dropped). A dropped invocation means a scheduled MCP health probe did not run — which is a monitoring blind spot, not just an infrastructure issue. The alarming strategy differs from typical AWS service alarms: you want InvocationDroppedCount > 0 with TreatMissingData: notBreaching (missing data means no invocations were attempted, which is fine) rather than a sum alarm with a fixed threshold. The second monitoring layer is the DLQ depth: the Scheduler DLQ captures invocations that could not reach the target at all (execution role failures, Lambda throttles, service unavailability), while Lambda's async destination captures function-level failures. Both queues need separate CloudWatch alarms. CloudTrail provides the audit trail for who created, modified, or deleted schedules — important for compliance and incident investigation when a probe disappears unexpectedly.

TL;DR

Create three alarms: (1) InvocationDroppedCount > 0 per schedule group with TreatMissingData: notBreaching — fires when any probe invocation is permanently dropped; (2) DLQ ApproximateNumberOfMessagesVisible > 0 with 1-minute period — fires when Scheduler cannot reach the target; (3) InvocationAttemptCount anomaly detection alarm to catch unexpected silences. Enable CloudTrail for the scheduler namespace to audit schedule changes. See the core scheduler guide for DLQ configuration and the groups guide for per-group metric dimensions.

CloudWatch metrics reference

All Scheduler metrics are in the AWS/Scheduler namespace. Metrics have two dimension sets: per-schedule (ScheduleName + ScheduleGroup) and per-group (ScheduleGroup only). Use per-group dimensions for aggregate alarming; use per-schedule dimensions for debugging individual probe failures.

MetricWhat it countsAlarm strategy
InvocationAttemptCountTotal invocation attempts (first + retries)Anomaly detection: alert if count drops to 0 for an extended period (schedule silently disabled or deleted)
InvocationDroppedCountInvocations dropped after exhausting retries OR MaximumEventAgeInSecondsCritical alarm: threshold > 0, TreatMissingData: notBreaching
InvocationThrottleCountInvocations throttled by Scheduler (not target throttling)Alert if persistent: Scheduler-level throttling indicates account-level rate limit hit
TargetErrorCountInvocations where the target returned an error (sync only)Alert for synchronous targets; for async Lambda targets this is 0 (async acceptance = success)
TargetErrorThrottledCountTarget throttle errors (e.g., Lambda TooManyRequestsException)Alert if persistent; indicates target capacity issue, not Scheduler issue
InvocationsSentToDeadLetterCountMessages sent to the configured DLQAlert on > 0 if you want DLQ-level alerting directly from metrics rather than from the SQS queue depth
# Query InvocationDroppedCount for a schedule group over the last hour
aws cloudwatch get-metric-statistics \
  --namespace AWS/Scheduler \
  --metric-name InvocationDroppedCount \
  --dimensions Name=ScheduleGroup,Value=mcp-monitoring-prod \
  --start-time "$(date -u -d '1 hour ago' '+%Y-%m-%dT%H:%M:%SZ')" \
  --end-time "$(date -u '+%Y-%m-%dT%H:%M:%SZ')" \
  --period 300 \
  --statistics Sum \
  --query 'Datapoints[*].{Time:Timestamp,Dropped:Sum}' \
  --output table

# Query per-schedule metrics (requires both dimensions)
aws cloudwatch get-metric-statistics \
  --namespace AWS/Scheduler \
  --metric-name InvocationAttemptCount \
  --dimensions \
    Name=ScheduleName,Value=mcp-health-probe-5m \
    Name=ScheduleGroup,Value=mcp-monitoring-prod \
  --period 300 \
  --statistics Sum
  # ... (start-time, end-time omitted for brevity)

Metrics are published with a delay of up to 5 minutes after the invocation. For rate(1 minute) schedules, you will see a 5-minute lag between the invocation and the metric appearing in CloudWatch. Design your alarm evaluation periods to accommodate this — use a 5-10 minute evaluation period rather than 1 minute to avoid false alarms from metric ingestion delay.

Critical alarm: InvocationDroppedCount

The most important Scheduler alarm is InvocationDroppedCount > 0. A dropped invocation is a permanently missed scheduled action — retries are exhausted, the event age limit passed, and Scheduler gave up without successfully invoking the target. For MCP monitoring, this is a monitoring gap: a probe that was supposed to run did not run, and you may have missed a real downtime event.

# CloudFormation / CDK CloudWatch alarm for InvocationDroppedCount
# (equivalent JSON for aws cloudwatch put-metric-alarm)
aws cloudwatch put-metric-alarm \
  --alarm-name "scheduler-invocation-dropped-mcp-prod" \
  --alarm-description "EventBridge Scheduler dropped invocations in mcp-monitoring-prod" \
  --namespace "AWS/Scheduler" \
  --metric-name "InvocationDroppedCount" \
  --dimensions Name=ScheduleGroup,Value=mcp-monitoring-prod \
  --period 300 \
  --evaluation-periods 1 \
  --threshold 0 \
  --comparison-operator GreaterThanThreshold \
  --statistic Sum \
  --treat-missing-data notBreaching \
  --alarm-actions "arn:aws:sns:us-east-1:123456789012:platform-alerts" \
  --ok-actions "arn:aws:sns:us-east-1:123456789012:platform-alerts"

# Key parameters explained:
# --threshold 0 + GreaterThanThreshold: alarm on ANY dropped invocation
# --treat-missing-data notBreaching: no data = no invocations attempted = fine
#   (do NOT use 'missing' or 'breaching' — those would alarm on quiet periods
#    between health probe intervals when no metric data points are published)
# --evaluation-periods 1: alarm immediately on first dropped invocation
#   (no need to wait for sustained failures — each drop is significant)

Use TreatMissingData: notBreaching (not missing or breaching). The AWS/Scheduler namespace only publishes data when invocations occur — between schedule intervals there are no data points. If you use TreatMissingData: breaching, the alarm fires between every probe interval (every 5 minutes for a rate(5 minutes) schedule) because there are no data points in the evaluation window. notBreaching correctly interprets the absence of data as "no invocations occurred in this period, which is expected."

DLQ monitoring

The Scheduler DLQ captures invocations that could not reach the target after all retries. Monitoring the DLQ queue depth provides a second alarm layer with different sensitivity than the CloudWatch metric alarm:

# Alarm on DLQ depth > 0 (any message = a failed invocation)
aws cloudwatch put-metric-alarm \
  --alarm-name "scheduler-dlq-depth-mcp-prod" \
  --alarm-description "Scheduler DLQ has messages — invocation failed to reach target" \
  --namespace "AWS/SQS" \
  --metric-name "ApproximateNumberOfMessagesVisible" \
  --dimensions Name=QueueName,Value=scheduler-dlq \
  --period 60 \
  --evaluation-periods 1 \
  --threshold 0 \
  --comparison-operator GreaterThanThreshold \
  --statistic Maximum \
  --treat-missing-data notBreaching \
  --alarm-actions "arn:aws:sns:us-east-1:123456789012:platform-alerts"

# DLQ message structure — inspect failed invocations
aws sqs receive-message \
  --queue-url "https://sqs.us-east-1.amazonaws.com/123456789012/scheduler-dlq" \
  --attribute-names All \
  --max-number-of-messages 10

# DLQ message body contains:
# {
#   "scheduleArn": "arn:aws:scheduler:us-east-1:123456789012:schedule/mcp-monitoring-prod/mcp-health-probe-5m",
#   "scheduledTime": "2026-10-04T14:00:00Z",
#   "attemptCount": 2,
#   "errorCode": "TargetLambdaPermissionError",
#   "errorMessage": "...",
#   "payload": "{...original input...}"
# }

The DLQ message body identifies the exact schedule, scheduled slot time, and error code. Common error codes in the DLQ:

Error code in DLQCauseResolution
TargetLambdaPermissionErrorExecution role lacks lambda:InvokeFunction permissionAdd InvokeFunction on the Lambda ARN to the execution role; check for alias/version mismatch
TargetThrottledTarget returned throttle (Lambda TooManyRequests, SQS rate limit)Increase Lambda reserved concurrency; add FlexibleTimeWindow to spread invocations
SchedulerRolePermissionErrorExecution role trust policy wrong or role deletedVerify trust policy names scheduler.amazonaws.com; recreate role if deleted
TargetFunctionErrorSynchronous Lambda invocation returned function errorCheck Lambda CloudWatch Logs for the function error; async invocations do not produce this code (async acceptance = success)

Detecting schedule silences with anomaly detection

A schedule that has been accidentally disabled or deleted produces zero metric data — not an elevated InvocationDroppedCount, but complete silence. Standard threshold alarms do not catch silence because TreatMissingData: notBreaching treats absence of data as fine. Use anomaly detection to catch unexpected drops in invocation volume:

# Anomaly detection alarm on InvocationAttemptCount
aws cloudwatch put-metric-alarm \
  --alarm-name "scheduler-invocation-silence-mcp-prod" \
  --alarm-description "Scheduler invocation rate dropped below expected — schedule may be disabled/deleted" \
  --metrics '[
    {
      "Id": "m1",
      "MetricStat": {
        "Metric": {
          "Namespace": "AWS/Scheduler",
          "MetricName": "InvocationAttemptCount",
          "Dimensions": [{"Name": "ScheduleGroup", "Value": "mcp-monitoring-prod"}]
        },
        "Period": 300,
        "Stat": "Sum"
      },
      "ReturnData": true
    },
    {
      "Id": "ad1",
      "Expression": "ANOMALY_DETECTION_BAND(m1, 2)",
      "ReturnData": true
    }
  ]' \
  --comparison-operator LessThanLowerOrGreaterThanUpperThreshold \
  --threshold-metric-id ad1 \
  --treat-missing-data notBreaching \
  --evaluation-periods 3 \
  --alarm-actions "arn:aws:sns:us-east-1:123456789012:platform-alerts"

# The anomaly detection band trains on historical invocation patterns.
# "2" is the standard deviation multiplier — increase to reduce false positives.
# evaluation-periods 3 = alarm only after 3 consecutive anomalous periods (15 minutes)
# to avoid alerting on brief dips from flexible time windows.

CloudTrail audit for Scheduler operations

EventBridge Scheduler operations are captured in CloudTrail as management events under the scheduler service. Auditing schedule changes is important for incident investigation — when a probe disappears, CloudTrail tells you who deleted it and when.

# Query CloudTrail for Scheduler events in the last 24 hours
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=EventSource,AttributeValue=scheduler.amazonaws.com \
  --start-time "$(date -u -d '24 hours ago' '+%Y-%m-%dT%H:%M:%SZ')" \
  --query 'Events[*].{Time:EventTime,Event:EventName,User:Username,Resources:Resources[*].ResourceName}' \
  --output table

# Common CloudTrail event names for Scheduler:
# scheduler:CreateSchedule      — who created a schedule and when
# scheduler:UpdateSchedule      — who changed a schedule (expression, state, target)
# scheduler:DeleteSchedule      — who deleted a schedule (KEY for incident investigation)
# scheduler:CreateScheduleGroup — who created a group
# scheduler:DeleteScheduleGroup — who deleted a group (cascade-deletes schedules)

# Look for suspicious DeleteSchedule events:
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=EventName,AttributeValue=DeleteSchedule \
  --start-time "$(date -u -d '7 days ago' '+%Y-%m-%dT%H:%M:%SZ')" \
  --query 'Events[*].{Time:EventTime,User:Username,Request:CloudTrailEvent}'

CloudTrail management events are enabled by default in most accounts. If you have CloudTrail disabled or are using a trail with management events filtered out, you will not have Scheduler change history. Enable a trail with IncludeManagementEvents: true before you need it for incident investigation — retroactive CloudTrail setup does not recover past events.

For automated compliance, use AWS Config with a custom rule or Security Hub finding rule to alert when schedules in a specific group are modified by anyone other than the authorized CI/CD role. CloudTrail + EventBridge Rules (not Scheduler) can trigger a Lambda that validates schedule changes against a policy and rolls back unauthorized modifications.

Schedule inventory and hygiene

# Full schedule inventory — all groups and states
for group in $(aws scheduler list-schedule-groups \
    --query 'ScheduleGroups[*].Name' --output text); do
  count=$(aws scheduler list-schedules \
    --group-name "$group" \
    --query 'length(Schedules)' --output text)
  enabled=$(aws scheduler list-schedules \
    --group-name "$group" \
    --query 'length(Schedules[?State==`ENABLED`])' --output text)
  disabled=$(aws scheduler list-schedules \
    --group-name "$group" \
    --query 'length(Schedules[?State==`DISABLED`])' --output text)
  completed=$(aws scheduler list-schedules \
    --group-name "$group" \
    --query 'length(Schedules[?State==`COMPLETED`])' --output text)
  echo "Group: $group | Total: $count | Enabled: $enabled | Disabled: $disabled | Completed: $completed"
done

# Find stale DISABLED schedules (disabled more than 30 days ago via tags or naming convention)
# Note: Scheduler does not track when a schedule was last modified via list-schedules
# — use CloudTrail to find UpdateSchedule events for DISABLED state changes

# Find schedules with no recent invocation (potential dead schedules)
# Check InvocationAttemptCount=0 over last 7 days per schedule
aws cloudwatch get-metric-data \
  --metric-data-queries '[{
    "Id": "invocations",
    "MetricStat": {
      "Metric": {
        "Namespace": "AWS/Scheduler",
        "MetricName": "InvocationAttemptCount",
        "Dimensions": [
          {"Name": "ScheduleGroup", "Value": "mcp-monitoring-prod"},
          {"Name": "ScheduleName", "Value": "mcp-health-probe-5m"}
        ]
      },
      "Period": 86400,
      "Stat": "Sum"
    },
    "ReturnData": true
  }]' \
  --start-time "$(date -u -d '7 days ago' '+%Y-%m-%dT%H:%M:%SZ')" \
  --end-time "$(date -u '+%Y-%m-%dT%H:%M:%SZ')" \
  --query 'MetricDataResults[0].Values'

Regular schedule hygiene prevents quota accumulation and reduces confusion when diagnosing probe failures. At minimum: (1) delete COMPLETED one-time schedules monthly; (2) review DISABLED schedules quarterly — if a schedule has been disabled for more than 30 days without a documented reason, it is likely safe to delete; (3) alert when InvocationAttemptCount for an ENABLED schedule is zero over 24 hours — this indicates a misconfigured cron expression or a silent Scheduler-level failure.

Enable/disable patterns for maintenance windows

# Pause all probes in a group for a maintenance window
# Step 1: list all ENABLED schedules
SCHEDULES=$(aws scheduler list-schedules \
  --group-name mcp-monitoring-prod \
  --query 'Schedules[?State==`ENABLED`].Name' \
  --output text)

# Step 2: disable each (requires full update-schedule with all current config)
for name in $SCHEDULES; do
  # Get current config
  config=$(aws scheduler get-schedule \
    --name "$name" \
    --group-name mcp-monitoring-prod)

  # Extract fields and resubmit with DISABLED state
  # (real implementation: use a script that parses and re-assembles the JSON)
  echo "Would disable: $name"
done

# Better pattern: use a separate "maintenance" schedule group
# During maintenance: move to maintenance mode by disabling the prod group's schedules
# After maintenance: re-enable

# Even better: use a single "maintenance-mode" flag in SSM Parameter Store
# Each Lambda probe checks the flag at start and exits immediately if set
# This avoids update-schedule calls entirely — no API calls needed for maintenance

The SSM Parameter Store pattern is simpler than disabling individual schedules for maintenance windows. Store a flag like /mcp-monitoring/maintenance-mode set to "true" during maintenance. Each Lambda probe function reads this parameter at startup (cache with 60-second TTL using SSM Parameter Store caching) and returns a 200 without probing if maintenance mode is active. This approach requires zero Scheduler API calls for maintenance windows — just one SSM put-parameter command — and the schedules keep firing (maintaining InvocationAttemptCount data continuity) while the probes themselves are safely no-oped.

Failure modes reference

FailureSymptomFix
InvocationDroppedCount alarm false-fires on quiet periodsAlarm ALARM state between probe intervalsSet TreatMissingData: notBreaching — Scheduler only emits metrics when invocations occur; absence of data between intervals is normal
DLQ alarm fires but DLQ empty after readingMessages consumed but alarm still ALARMSQS ApproximateNumberOfMessagesVisible lags after consume; wait for alarm to auto-resolve; add a DLQ consumer that logs message content before deleting
CloudTrail shows DeleteScheduleGroup but no individual DeleteSchedule eventsCascade-delete: all schedules deleted but only group deletion visibleExpected: cascade-delete does not emit individual DeleteSchedule events; CloudTrail only shows DeleteScheduleGroup; document cascade-delete risk in runbooks
Anomaly detection alarm fires on initial training periodAlarm INSUFFICIENT_DATA or ALARM in first 2 weeksAnomaly detection requires ~2 weeks of historical data to train; suppress or ignore alarm actions during the training period; use threshold alarm as backup
InvocationAttemptCount=0 for ENABLED scheduleNo invocations visible in metrics despite ENABLED stateCheck cron expression validity — an invalid cron may have been accepted at creation but never fires; test with a rate() expression; verify AWS/Scheduler namespace in correct region