AWS EventBridge Scheduler · 2026-10-08 · EventBridge Scheduler arc

AWS EventBridge Scheduler for MCP Servers: Schedule Expressions, SDK Targets, and Production Observability

The single most disruptive mistake in an EventBridge Scheduler setup is using events.amazonaws.com instead of scheduler.amazonaws.com as the trust principal in the execution role — the schedule creates successfully, the API returns no error, but every invocation attempt silently fails because AWS refuses to assume a role whose trust policy doesn't include the Scheduler service principal, and the only symptom is an empty InvocationAttemptCount metric with no corresponding alarm unless you've pre-configured InvocationDroppedCount monitoring. For MCP server health probes, EventBridge Scheduler solves the "invoke a probe every 5 minutes regardless of what's running" problem without managing cron workers or Lambda triggers — a rate expression fires a Lambda or ECS task on a fixed interval, the FlexibleTimeWindow setting spreads load across hundreds of concurrent schedules to prevent cold start pile-ups, and the retry policy with MaximumEventAgeInSeconds: 300 discards stale probe invocations rather than re-probing a slot that closed 20 minutes ago. The five topics that produce a production-ready scheduled probe system — schedule expressions and execution role setup (Scheduler vs Rules vs Pipes distinction; scheduler.amazonaws.com principal with SourceArn condition; rate/cron/at expressions; FlexibleTimeWindow; retry policy and DLQ; five common failure modes), SDK targets for direct AWS API invocation (270+ SDK targets without Lambda intermediary; ARN format arn:aws:scheduler:::aws-sdk:service:action; ECS RunTask, Step Functions StartExecution, SQS SendMessage, DynamoDB PutItem patterns; <$.context.scheduledTime> and other context variables; idempotency strategies per target type), Lambda targets and the double-retry problem (async vs sync invocation; the 6× invocation storm from combined Scheduler retries + Lambda internal retries; disabling Lambda's EventInvokeConfig retries; idempotency with scheduledTime composite key; two-layer DLQ architecture; cross-account Lambda targets), schedule groups and quota management (default vs named groups; cascade-delete danger; quota hierarchy 10,000/group; one-time schedule cleanup; per-group IAM scoping; IaC with Terraform and CDK), and CloudWatch metrics and operational visibility (six metrics in AWS/Scheduler namespace; InvocationDroppedCount as the critical alarm; DLQ error codes; CloudTrail audit events; anomaly detection for silent schedule failures; maintenance window pattern with SSM Parameter Store) — each contains sharp edges that produce silent probe failures, runaway retry storms, or quota exhaustion that only surfaces when a maintenance window is already open. The execution role iam:PassRole must be scoped to the specific target's role ARN — an unrestricted PassRole passes an AWS Security Hub finding and allows privilege escalation. update-schedule requires ALL fields to be re-specified: omitting any optional field silently resets it to default, so a partial update that only changes ScheduleExpression will also reset MaximumRetryAttempts to 185 and MaximumEventAgeInSeconds to 86400 unless both are included. This guide synthesizes all five topics into three structural patterns: schedule configuration and the trust principal gap, target integration patterns (SDK targets and Lambda), and groups, quotas, and observability.

TL;DR

Pattern 1 — Schedule Configuration and the Trust Principal Gap

The scheduler.amazonaws.com principal and why it matters

EventBridge has three distinct invocation mechanisms that are easy to conflate: EventBridge Rules (event pattern matching on an event bus, using the events.amazonaws.com service principal), EventBridge Pipes (streaming source transformations for SQS/Kinesis/DynamoDB Streams, also on the event bus), and EventBridge Scheduler (time-based scheduled invocations, completely independent of the event bus). The Scheduler service uses its own IAM service principal: scheduler.amazonaws.com. Using the wrong principal — the one copied from a EventBridge Rules setup — is the most common first-deployment failure mode because the schedule creates without error, the IAM role exists and has the right action permissions, but the Scheduler service cannot assume the role because the trust policy excludes it. The invocation attempts happen on schedule, hit the role assumption failure, and are silently dropped. The only observable symptom is that InvocationAttemptCount stays at zero — but if you haven't pre-configured that alarm, you won't notice until your probe SLA misses.

The correct execution role trust policy for EventBridge Scheduler:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "Service": "scheduler.amazonaws.com" },
    "Action": "sts:AssumeRole",
    "Condition": {
      "StringEquals": {
        "aws:SourceAccount": "123456789012"
      },
      "ArnLike": {
        "aws:SourceArn": "arn:aws:scheduler:us-east-1:123456789012:schedule/*/*"
      }
    }
  }]
}

The SourceArn wildcard schedule/*/* covers all groups and schedule names in the account-region — this is the recommended scope for a shared execution role. The SourceAccount condition prevents confused deputy attacks where another account's Scheduler could assume this role. For per-group role isolation, narrow SourceArn to schedule/mcp-monitoring-prod/*.

Rate, cron, and one-time expressions

EventBridge Scheduler supports three expression types. Rate expressions use rate(value unit) where unit is minute, minutes, hour, hours, day, or days — singular for value 1, plural otherwise (rate(1 minute), rate(5 minutes)). Rate expressions start immediately on schedule creation and fire continuously — they're the simplest and most predictable form for regular MCP health probes.

Cron expressions use six fields: cron(minutes hours day-of-month month day-of-week year). Three critical differences from standard Unix cron: (1) there is no seconds field — the scheduler fires at the start of the specified minute; (2) the year field is required; (3) either day-of-month or day-of-week must be ? (but not both, and not neither) — specifying both as * produces a ValidationException. Example for 6 AM UTC daily: cron(0 6 * * ? *). For the first Monday of each month: cron(0 9 ? * 2#1 *).

One-time expressions use at(yyyy-mm-ddThh:mm:ss) — always in UTC, no timezone suffix. One-time schedules auto-complete after firing and transition to COMPLETED state. Always set ActionAfterCompletion: DELETE when creating one-time schedules: completed schedules count toward the 10,000/group quota and do not self-delete by default. A deployment that creates thousands of maintenance window probes without this flag will exhaust the group quota within months.

FlexibleTimeWindow and retry policy

FlexibleTimeWindow controls invocation timing within the scheduled window. With FlexibleTimeWindowMode: OFF (default), all schedules with the same interval fire simultaneously — at minute 0, all rate(5 minute) schedules across the group fire at exactly the same second, causing Lambda cold starts and function concurrency spikes. With FlexibleTimeWindowMode: FLEXIBLE and MaximumWindowInMinutes: 5, the Scheduler distributes invocations randomly within the 5-minute window. Enable this for any group with more than 50 concurrent schedules firing on the same interval.

Retry policy has two parameters: MaximumRetryAttempts (0–185, default 185) and MaximumEventAgeInSeconds (60–86400, default 86400). For MCP health probes, set MaximumEventAgeInSeconds: 300 — an invocation that failed to fire more than 5 minutes ago represents stale probe data, and retrying it at minute 7 produces a false "ok" result for a window that was actually unchecked. For ECS RunTask targets, set MaximumRetryAttempts: 0 — retrying a RunTask starts a second container instance, and the first container (which may have succeeded or partially written) remains running in parallel.

The DLQ captures invocation-level failures after retry exhaustion. The execution role needs sqs:SendMessage on the DLQ queue ARN, and the SQS queue resource policy must allow the Scheduler service principal to write:

{
  "Effect": "Allow",
  "Principal": { "Service": "scheduler.amazonaws.com" },
  "Action": "sqs:SendMessage",
  "Resource": "arn:aws:sqs:us-east-1:123456789012:mcp-scheduler-dlq",
  "Condition": {
    "ArnEquals": { "aws:SourceArn": "arn:aws:scheduler:us-east-1:123456789012:schedule/*/*" }
  }
}

Missing the SQS queue resource policy is a silent failure: the Scheduler cannot write to the DLQ, so dropped invocations are lost with no observable record. The only way to detect missing DLQ permissions is to notice that InvocationDroppedCount is positive while InvocationsSentToDeadLetterCount stays at zero.

The update-schedule full-replacement trap

Unlike most AWS API update operations, update-schedule is a full replacement: any field not included in the call is reset to its service default. If you call update-schedule specifying only a new ScheduleExpression, the call succeeds and the expression changes — but MaximumRetryAttempts resets to 185, MaximumEventAgeInSeconds resets to 86400, DeadLetterConfig is cleared, and FlexibleTimeWindow reverts to OFF. Always retrieve the current schedule with get-schedule before updating and include all fields in the update call. Infrastructure-as-code tools that manage schedules declaratively (Terraform, CDK) handle this automatically — the risk is in manual CLI operations during incident response.

Pattern 2 — Target Integration Patterns: SDK Targets and Lambda

SDK targets and the direct API invocation model

EventBridge Scheduler can invoke over 270 AWS SDK APIs directly without a Lambda intermediary — the invocation passes a JSON input directly to the service API, which eliminates cold start latency and Lambda cost for simple scheduled operations. The target ARN format is arn:aws:scheduler:::aws-sdk:<service>:<camelCaseAction> — note the triple colon (no region or account in the ARN, because the target is the global SDK operation, not a specific resource). The input must be a JSON string matching the SDK API's request parameter structure exactly.

Context variables let you inject invocation metadata into the target input. These are template strings that the Scheduler evaluates at invocation time:

Context variables use angle-bracket syntax (<$.context.scheduledTime>), not the EventBridge input transformer dollar-sign syntax. Incorrect syntax causes the variable to be passed as a literal string — the schedule creates and fires, but the downstream function receives the raw template text instead of the timestamp value.

ECS RunTask: scheduling container-based probes

The ECS RunTask SDK target (arn:aws:scheduler:::aws-sdk:ecs:runTask) is well-suited for MCP endpoint health probes where the probe logic is packaged as a Fargate container — 30-60 second tasks that probe an MCP endpoint, write results to DynamoDB, and exit. The execution role requires three IAM permissions that must all be present: ecs:RunTask on the task definition ARN, iam:PassRole on both the ECS task execution role ARN and the task role ARN (missing either PassRole triggers an AccessDenied that superficially resembles an ECS permission error), and sqs:SendMessage on the DLQ queue ARN. Pass <$.context.scheduledTime> as a container environment variable for idempotency — the task can write a DynamoDB record keyed on the scheduled time to detect and skip duplicate runs.

Set MaximumRetryAttempts: 0 for ECS RunTask targets. When Scheduler retries a RunTask that previously returned an error, it starts a second container — the first container may still be running, partially writing results. Two concurrent probe containers writing to the same DynamoDB partition key produces interleaved results. The DLQ captures the failed invocation attempt for manual investigation; the retry cost is borne by running duplicate containers in production.

Step Functions StartExecution: idempotent workflow scheduling

For MCP server workflows that involve multiple steps — probe, analyze, alert, record — Step Functions StartExecution as a Scheduler target is cleaner than a Lambda that calls Step Functions. The execution name must be unique within 90 days; use <$.context.scheduledTime> to construct it: "hourly-mcp-probe-<$.context.scheduledTime>". If Scheduler retries and the execution name already exists, Step Functions returns ExecutionAlreadyExists — this is non-fatal from the workflow perspective, but Scheduler treats it as a target error and may continue retrying.

Use Express Workflows for sub-hourly schedules: Express Workflows support up to 5-minute duration, cost per invocation (not per state transition), and are suitable for high-frequency short-lived probe workflows. Standard Workflows are appropriate for hourly or less frequent probes where duration can exceed 5 minutes or where the execution history audit trail is required.

SQS and DynamoDB SDK targets

The SQS SendMessage target (arn:aws:scheduler:::aws-sdk:sqs:sendMessage) queues a probe trigger for downstream processing — the execution role needs sqs:SendMessage on the queue ARN and sqs:GetQueueAttributes for FIFO queues. For FIFO queues, the input must include both MessageGroupId and MessageDeduplicationId — use <$.context.scheduledTime> as the deduplication ID (5-minute deduplication window matches the typical probe interval). Note that the Scheduler's DLQ and the SQS queue's own DLQ are distinct: the Scheduler DLQ captures invocation-level failures (Scheduler couldn't send the message), while the queue's DLQ captures processing failures after the message is received.

The DynamoDB PutItem target (arn:aws:scheduler:::aws-sdk:dynamodb:putItem) requires DynamoDB's typed JSON format — {"S":"value"} for strings, {"N":"42"} for numbers, not plain JSON values. Passing a plain string causes a SerializationException at invocation time. Add a ConditionExpression: "attribute_not_exists(sk)" to the input for at-most-once writes — a ConditionalCheckFailedException on retry indicates the record already exists, which is the correct behavior. Without this condition, a Scheduler retry overwrites the original probe record with stale data.

Lambda targets and the double-retry problem

Lambda targets work differently depending on invocation type. Async invocation (default InvocationType: Event): Scheduler fires the invocation and receives HTTP 202 (accepted) from Lambda — this is immediately counted as a successful invocation from Scheduler's perspective, regardless of whether the function executes successfully. Function errors go to Lambda's async destination (EventInvokeConfig.DestinationConfig.OnFailure), not the Scheduler DLQ. This asymmetry is the source of the double-retry problem.

Lambda async invocations have a built-in internal retry configured in EventInvokeConfig: the default is 2 additional attempts after the first failure, with exponential backoff over up to 6 hours. Combined with Scheduler's default of 185 retries, a single throttle produces a compounding retry cascade: Scheduler fires the invocation (attempt 1), Lambda accepts it (HTTP 202), Lambda's internal retries add 2 more attempts (attempts 2 and 3), Lambda sends to Lambda's async destination after exhaustion. Meanwhile, because Scheduler received HTTP 202 it considers the invocation successful and does not retry. The net result is 3 Lambda invocations per scheduled firing — not 6 — but if the probe is hitting a throttled endpoint, those 3 invocations all probe within a 6-hour window, poisoning the probe result timeline.

For health probes, disable Lambda's internal retry entirely:

aws lambda put-function-event-invoke-config \
  --function-name mcp-health-probe \
  --maximum-retry-attempts 0 \
  --maximum-event-age-in-seconds 300

This means exactly one Lambda invocation per Scheduler firing. If the function throws, the invocation is lost (no internal retry). The Scheduler DLQ is unaffected — it only receives messages for invocation-level failures (Lambda throttle, wrong principal, timeout before 202 acceptance), not function execution failures. Configure Lambda's async destination for function-level failures separately.

Idempotency with scheduledTime as composite key

The <$.context.scheduledTime> variable provides the nominal scheduled firing time — this value is the same across all retry attempts for a given slot, making it suitable as an idempotency key. A conditional DynamoDB write guards against duplicate processing:

import boto3, os, json
from botocore.exceptions import ClientError

dynamodb = boto3.resource('dynamodb')
table = dynamodb.Table(os.environ['RESULTS_TABLE'])

def handler(event, context):
    scheduled_time = event.get('scheduledTime')
    schedule_arn = event.get('scheduleArn')
    # Composite key: scheduleArn + scheduledTime
    pk = f"{schedule_arn}#{scheduled_time}"
    try:
        table.put_item(
            Item={'pk': pk, 'scheduled_time': scheduled_time, 'status': 'processing'},
            ConditionExpression='attribute_not_exists(pk)'
        )
    except ClientError as e:
        if e.response['Error']['Code'] == 'ConditionalCheckFailedException':
            return {'statusCode': 200, 'body': 'duplicate, skipped'}
        raise
    # ... run probe ...

Use the composite key scheduleArn + scheduledTime rather than scheduledTime alone — multiple schedules may fire at the same UTC second in a large multi-tenant deployment, and scheduledTime alone would prevent the second schedule's invocation from writing its result. Log attemptNumber for debugging, but do not include it in the idempotency key.

Cross-account Lambda targets

When the Scheduler is in an operations account and the Lambda probe function is in a target account, two IAM configurations are required: a Lambda resource policy in the target account granting the scheduler.amazonaws.com service principal lambda:InvokeFunction with conditions on aws:SourceAccount and aws:SourceArn scoped to the ops account's schedule ARN pattern; and an execution role in the ops account with lambda:InvokeFunction on the target account's Lambda ARN. The execution role trust policy uses scheduler.amazonaws.com as normal. A common mistake is specifying the execution role ARN (not the service principal) in the Lambda resource policy — the Lambda service validates the calling principal as scheduler.amazonaws.com, not the assumed-role ARN.

Pattern 3 — Groups, Quotas, and Observability

Schedule group architecture for multi-tenant MCP monitoring

The default schedule group (default) is created automatically, cannot be deleted or renamed, and accepts schedules when no --group-name is specified. For production MCP monitoring, create named groups per environment: mcp-monitoring-prod, mcp-monitoring-staging, mcp-monitoring-dev. Named groups enable per-environment IAM scoping, cost allocation via tags, and group-level quota visibility.

Tag groups at creation for cost allocation — tags cannot be set after creation via create-schedule-group in some SDK versions; add them immediately with tag-resource:

aws scheduler create-schedule-group \
  --name mcp-monitoring-prod \
  --region us-east-1

aws scheduler tag-resource \
  --resource-arn arn:aws:scheduler:us-east-1:123456789012:schedule-group/mcp-monitoring-prod \
  --tags Environment=prod,Team=platform,CostCenter=eng-infra,Service=mcp-uptime

Cascade-delete and IaC protection

delete-schedule-group deletes all schedules in the group atomically and irreversibly. There is no --dry-run flag, no confirmation prompt, and no recovery path. CloudTrail records the DeleteScheduleGroup event but does NOT emit individual DeleteSchedule events for each deleted schedule — after a cascade-delete, you cannot reconstruct the deleted schedule list from CloudTrail alone.

Protect production groups in infrastructure-as-code before any schedules are created in them:

# Terraform
resource "aws_scheduler_schedule_group" "prod" {
  name = "mcp-monitoring-prod"
  lifecycle {
    prevent_destroy = true
  }
}

For CDK, use a CloudFormation deletion policy on the CfnScheduleGroup L1 construct. If a production group is accidentally deleted during an IaC teardown, you must recreate the group, recreate all schedules (from IaC state or runbooks), and accept a monitoring gap equal to the recreation time.

Quota hierarchy and scaling limits

Quota Default Limit Adjustable Notes
Schedules per account-region 1,000,000 No Hard ceiling; cannot be raised via Service Quotas
Schedules per group 10,000 No Applies to default group and every named group
Schedule groups per account-region 500 Yes Request increase via Service Quotas
CreateSchedule / UpdateSchedule / DeleteSchedule TPS 10 Yes ListSchedules / GetSchedule: 100 TPS

At 500 groups × 10,000 schedules = 5,000,000 theoretical capacity, but the account ceiling of 1,000,000 is the binding constraint. For large multi-tenant deployments exceeding 200,000 schedules, distribute across multiple AWS accounts rather than increasing group count — the account ceiling is unadjustable.

One-time schedule cleanup prevents group quota exhaustion. Completed schedules persist as COMPLETED state and count toward the 10,000/group limit. List and delete them monthly:

aws scheduler list-schedules \
  --group-name mcp-monitoring-prod \
  --state COMPLETED \
  --query 'Schedules[].Name' \
  --output text | tr '\t' '\n' | while read name; do
    aws scheduler delete-schedule --name "$name" --group-name mcp-monitoring-prod
done

Per-group IAM scoping

Scope IAM actions to specific groups using the resource ARN pattern:

{
  "Effect": "Allow",
  "Action": [
    "scheduler:CreateSchedule",
    "scheduler:UpdateSchedule",
    "scheduler:DeleteSchedule",
    "scheduler:GetSchedule",
    "scheduler:ListSchedules"
  ],
  "Resource": "arn:aws:scheduler:us-east-1:123456789012:schedule/mcp-monitoring-prod/*"
}

Without group scoping, a developer role with scheduler:* can modify schedules in any group including production. The schedule ARN format is arn:aws:scheduler:region:account:schedule/GROUP-NAME/SCHEDULE-NAME — the group name is part of the ARN, enabling resource-level permissions without tag-based conditions.

CloudWatch metrics for scheduler observability

EventBridge Scheduler publishes metrics in the AWS/Scheduler namespace with up to 5-minute delay. Six metrics are available:

Metric What It Counts Alarm Strategy
InvocationAttemptCount Total invocation attempts (first + retries) Anomaly detection: alert if drops to 0 unexpectedly
InvocationDroppedCount Dropped after retry exhaustion or event age exceeded CRITICAL: > 0, TreatMissingData: notBreaching
InvocationThrottleCount Scheduler-level throttling (API rate limit) Alert if persistent (rate limit hit)
TargetErrorCount Target returned error (sync targets only) Alert for sync targets; always 0 for async Lambda
TargetErrorThrottledCount Target throttle errors Alert if persistent (target capacity problem)
InvocationsSentToDeadLetterCount Messages successfully sent to DLQ Alert on > 0 for direct DLQ monitoring

The InvocationDroppedCount alarm is the highest-priority operational signal — a non-zero value means probe invocations are being permanently lost. The alarm configuration requires special care with TreatMissingData:

aws cloudwatch put-metric-alarm \
  --alarm-name "mcp-scheduler-dropped-invocations" \
  --namespace AWS/Scheduler \
  --metric-name InvocationDroppedCount \
  --dimensions Name=ScheduleGroup,Value=mcp-monitoring-prod \
  --statistic Sum \
  --period 300 \
  --threshold 0 \
  --comparison-operator GreaterThanThreshold \
  --evaluation-periods 1 \
  --treat-missing-data notBreaching \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-ops-alerts

Use TreatMissingData: notBreaching. Using missing holds the alarm in its previous state (ambiguous during quiet periods). Using breaching alarms whenever no probe has fired in the evaluation period — which is expected behavior for schedules that fire every 5 minutes within a 1-minute evaluation window. The metric only appears when an invocation fires, so gaps between probe intervals produce missing data points; notBreaching treats those gaps as normal.

Anomaly detection for silenced schedules

A schedule that was accidentally disabled or deleted shows zero InvocationAttemptCount — but a fixed-threshold alarm on InvocationAttemptCount = 0 fires during normal quiet periods. Use CloudWatch anomaly detection instead:

aws cloudwatch put-metric-alarm \
  --alarm-name "mcp-scheduler-silent-probe-anomaly" \
  --namespace AWS/Scheduler \
  --metric-name InvocationAttemptCount \
  --dimensions Name=ScheduleGroup,Value=mcp-monitoring-prod \
  --comparison-operator LessThanLowerOrGreaterThanUpperThreshold \
  --evaluation-periods 3 \
  --datapoints-to-alarm 3 \
  --treat-missing-data notBreaching \
  --metrics '[{"Id":"m1","MetricStat":{"Metric":{"Namespace":"AWS/Scheduler","MetricName":"InvocationAttemptCount","Dimensions":[{"Name":"ScheduleGroup","Value":"mcp-monitoring-prod"}]},"Period":300,"Stat":"Sum"}},{"Id":"e1","Expression":"ANOMALY_DETECTION_BAND(m1, 2)","Label":"InvocationAttemptCount (expected)"}]' \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-ops-alerts

Anomaly detection requires approximately 2 weeks of training data before it models the baseline accurately — during the training period, suppress the alarm or accept false positives. After training, it reliably catches schedules that fire at 0 when the trained baseline predicts a positive count.

DLQ error codes and resolution patterns

Each DLQ message contains a JSON body with scheduleArn, scheduledTime, attemptCount, errorCode, errorMessage, and payload. The most common error codes:

CloudTrail audit and the cascade-delete gap

Enable CloudTrail management events in the account-region before any production Scheduler usage — there is no retroactive capture. Scheduler API calls appear under service name scheduler.amazonaws.com. Key events to monitor: CreateSchedule, UpdateSchedule, DeleteSchedule, CreateScheduleGroup, DeleteScheduleGroup. The cascade-delete gap: when DeleteScheduleGroup fires, CloudTrail records only the group deletion event — it does not emit individual DeleteSchedule events for the hundreds or thousands of schedules deleted in the cascade. After a cascade-delete incident, the only record is the single DeleteScheduleGroup event with the deleting principal's identity; the deleted schedule names are not recoverable from CloudTrail. Maintain IaC state as the canonical source of schedule definitions, not CloudTrail.

Maintenance window pattern with SSM Parameter Store

During infrastructure maintenance, disabling hundreds of Scheduler schedules individually via update-schedule --state DISABLED is slow (10 TPS API limit), noisy in CloudTrail, and risky (the re-enable step may partially fail, leaving some schedules permanently disabled). A cleaner approach uses SSM Parameter Store as a maintenance flag:

import boto3, os, time

ssm = boto3.client('ssm')
_maintenance_cache = {'value': False, 'expires': 0}

def is_maintenance_mode():
    now = time.monotonic()
    if now < _maintenance_cache['expires']:
        return _maintenance_cache['value']
    try:
        resp = ssm.get_parameter(Name='/mcp-monitoring/maintenance-mode')
        val = resp['Parameter']['Value'].lower() == 'true'
    except ssm.exceptions.ParameterNotFound:
        val = False
    _maintenance_cache.update({'value': val, 'expires': now + 60})
    return val

def handler(event, context):
    if is_maintenance_mode():
        return {'statusCode': 200, 'body': 'maintenance mode, skipping probe'}
    # ... run probe ...

Schedules keep firing during maintenance — InvocationAttemptCount continues to increment, maintaining metric continuity and preventing anomaly detection alarms from triggering on a sudden drop to zero. Probes no-op immediately after the SSM check. To start maintenance: aws ssm put-parameter --name /mcp-monitoring/maintenance-mode --value true --overwrite. To end it: aws ssm put-parameter --name /mcp-monitoring/maintenance-mode --value false --overwrite. The 60-second TTL cache means probes transition out of maintenance mode within one minute of the parameter update — no update-schedule calls required.

Consolidated failure modes

# Failure Symptom Root cause Fix
1 Schedules create but never fire InvocationAttemptCount = 0 Execution role trust policy uses events.amazonaws.com not scheduler.amazonaws.com Update trust policy principal
2 ValidationException on cron creation API error at create-schedule Both day-of-month and day-of-week are * — one must be ? Set ? in exactly one of the two fields
3 Group quota exhaustion despite room CreateSchedule fails with quota error COMPLETED one-time schedules not cleaned up; count toward 10,000/group limit Delete COMPLETED schedules; set ActionAfterCompletion: DELETE
4 Partial update resets retry policy MaximumRetryAttempts reverts to 185 after update update-schedule is full replacement; omitted fields reset to defaults Always include all fields in update-schedule calls
5 ECS duplicate container instances Two probe containers running concurrently MaximumRetryAttempts > 0 on RunTask target; Scheduler retries start second container Set MaximumRetryAttempts: 0 for ECS RunTask targets
6 SDK target ValidationException API error at schedule creation SDK action not supported, or camelCase capitalization wrong in ARN Verify ARN against supported targets list; check capitalization
7 Context variables passed as literal strings Function receives <$.context.scheduledTime> as text Incorrect syntax (using ${scheduledTime} instead of <$.context.scheduledTime>) Use angle-bracket syntax: <$.context.scheduledTime>
8 Lambda double-retry storm 3+ Lambda invocations per scheduled firing Lambda EventInvokeConfig MaximumRetryAttempts default = 2 Set maximum-retry-attempts 0 via put-function-event-invoke-config
9 Async Lambda errors not in Scheduler DLQ Scheduler DLQ empty despite function failures Async invocation: HTTP 202 = Scheduler success; function errors go to Lambda destination Configure Lambda EventInvokeConfig OnFailure destination separately
10 DynamoDB typed JSON error SerializationException at invocation time DynamoDB SDK target requires typed JSON format ({"S":"value"}) Use DynamoDB typed format in Input JSON
11 DLQ permission missing — dropped invocations lost InvocationDroppedCount > 0, InvocationsSentToDeadLetterCount = 0 SQS queue resource policy missing scheduler.amazonaws.com principal Add SQS resource policy allowing scheduler.amazonaws.com SendMessage
12 Cascade-delete destroys production schedules All schedules in group deleted delete-schedule-group deletes all schedules atomically Add lifecycle prevent_destroy in Terraform; use CloudFormation deletion protection
13 InvocationDroppedCount alarm false-fires Alarm fires during quiet periods with no dropped invocations TreatMissingData set to missing or breaching Set TreatMissingData: notBreaching
14 Anomaly detection alarm fires during training period False positive alarms in first 2 weeks Model training period ~14 days; baseline not yet established Suppress anomaly detection alarm for first 2 weeks after enabling
15 Step Functions ExecutionAlreadyExists causes retry loop Scheduler retries indefinitely for a succeeded workflow ExecutionAlreadyExists is treated as a target error; Scheduler retries the invocation Use Express Workflows (no execution name uniqueness requirement); or set MaximumRetryAttempts: 0

Production checklists

Execution role and schedule creation

Target integration

Groups and operations

Further reading