AWS EventBridge Scheduler · 2026-10-08 · EventBridge Scheduler arc
AWS EventBridge Scheduler for MCP Servers: Schedule Expressions, SDK Targets, and Production Observability
The single most disruptive mistake in an EventBridge Scheduler setup is using events.amazonaws.com instead of scheduler.amazonaws.com as the trust principal in the execution role — the schedule creates successfully, the API returns no error, but every invocation attempt silently fails because AWS refuses to assume a role whose trust policy doesn't include the Scheduler service principal, and the only symptom is an empty InvocationAttemptCount metric with no corresponding alarm unless you've pre-configured InvocationDroppedCount monitoring. For MCP server health probes, EventBridge Scheduler solves the "invoke a probe every 5 minutes regardless of what's running" problem without managing cron workers or Lambda triggers — a rate expression fires a Lambda or ECS task on a fixed interval, the FlexibleTimeWindow setting spreads load across hundreds of concurrent schedules to prevent cold start pile-ups, and the retry policy with MaximumEventAgeInSeconds: 300 discards stale probe invocations rather than re-probing a slot that closed 20 minutes ago. The five topics that produce a production-ready scheduled probe system — schedule expressions and execution role setup (Scheduler vs Rules vs Pipes distinction; scheduler.amazonaws.com principal with SourceArn condition; rate/cron/at expressions; FlexibleTimeWindow; retry policy and DLQ; five common failure modes), SDK targets for direct AWS API invocation (270+ SDK targets without Lambda intermediary; ARN format arn:aws:scheduler:::aws-sdk:service:action; ECS RunTask, Step Functions StartExecution, SQS SendMessage, DynamoDB PutItem patterns; <$.context.scheduledTime> and other context variables; idempotency strategies per target type), Lambda targets and the double-retry problem (async vs sync invocation; the 6× invocation storm from combined Scheduler retries + Lambda internal retries; disabling Lambda's EventInvokeConfig retries; idempotency with scheduledTime composite key; two-layer DLQ architecture; cross-account Lambda targets), schedule groups and quota management (default vs named groups; cascade-delete danger; quota hierarchy 10,000/group; one-time schedule cleanup; per-group IAM scoping; IaC with Terraform and CDK), and CloudWatch metrics and operational visibility (six metrics in AWS/Scheduler namespace; InvocationDroppedCount as the critical alarm; DLQ error codes; CloudTrail audit events; anomaly detection for silent schedule failures; maintenance window pattern with SSM Parameter Store) — each contains sharp edges that produce silent probe failures, runaway retry storms, or quota exhaustion that only surfaces when a maintenance window is already open. The execution role iam:PassRole must be scoped to the specific target's role ARN — an unrestricted PassRole passes an AWS Security Hub finding and allows privilege escalation. update-schedule requires ALL fields to be re-specified: omitting any optional field silently resets it to default, so a partial update that only changes ScheduleExpression will also reset MaximumRetryAttempts to 185 and MaximumEventAgeInSeconds to 86400 unless both are included. This guide synthesizes all five topics into three structural patterns: schedule configuration and the trust principal gap, target integration patterns (SDK targets and Lambda), and groups, quotas, and observability.
TL;DR
- Schedule configuration and the trust principal gap: create an EventBridge Scheduler execution role with trust principal
scheduler.amazonaws.com— notevents.amazonaws.com(EventBridge Rules) and notlambda.amazonaws.com. Scope the trust condition toSourceArn: "arn:aws:scheduler:region:account:schedule/*/*"(wildcard covers all groups and schedules in the account-region) and add aSourceAccountcondition to prevent confused deputy attacks. For schedule expressions: userate(5 minutes)for fixed-interval MCP health probes — the simplest and most reliable form; usecron(0 6 * * ? *)for time-of-day work (six fields, year required,?in exactly one of day-of-month or day-of-week, no seconds field); useat(2026-12-01T02:00:00)for one-time maintenance window probes — always setActionAfterCompletion: DELETEon one-time schedules or they accumulate as COMPLETED state and count toward the 10,000/group quota. SetFlexibleTimeWindowwithMaximumWindowInMinutes: 5for any deployment with more than 50 concurrent schedules — without it, all schedules on the same interval fire simultaneously and cause Lambda cold start pile-ups. For retry policy: setMaximumEventAgeInSeconds: 300for health probes (discard invocations more than 5 minutes old — stale probe data is worse than a gap); setMaximumRetryAttempts: 0for ECS RunTask targets (retrying a RunTask starts a second task); configure a DLQ as an SQS queue — the execution role needssqs:SendMessageon the queue ARN, and the queue resource policy must allow the Scheduler service principal to publish. Everyupdate-schedulecall must re-specify all fields: retry policy, DLQ, FlexibleTimeWindow, and State are silently reset to defaults if omitted. - Target integration patterns — SDK targets and Lambda: SDK targets let EventBridge Scheduler call 270+ AWS APIs directly without a Lambda intermediary — the ARN format is
arn:aws:scheduler:::aws-sdk:ecs:runTask(note triple colon: no region or account in the ARN). The Input field is a JSON string matching the SDK API's request structure, with context variables injected using<$.context.scheduledTime>,<$.context.scheduleArn>,<$.context.scheduleName>,<$.context.scheduleGroupName>, and<$.context.attemptNumber>(zero-based). For ECS RunTask: setMaximumRetryAttempts: 0— retrying a failed RunTask starts a second container; the execution role needsecs:RunTaskon the task definition ARN,iam:PassRoleon both the ECS task execution role and task role ARNs (missing either causes AccessDenied that looks like an ECS permission problem), andsqs:SendMessagefor the DLQ. For Step Functions StartExecution: use<$.context.scheduledTime>in the execution name to guarantee uniqueness within 90 days — if a retry occurs and the execution already exists, theExecutionAlreadyExistsexception is non-fatal but Scheduler treats it as a failure and retries; use Express Workflows (5-minute max, much lower cost) for sub-hourly schedules. For DynamoDB PutItem: use DynamoDB's typed JSON format ({"S":"value"}not plain strings) and addConditionExpression: "attribute_not_exists(sk)"for at-most-once idempotency — aConditionalCheckFailedExceptionon retry is expected and handled. For Lambda targets: async invocation (default) means Scheduler success equals Lambda accepting the invocation (HTTP 202), regardless of function execution outcome — function errors land in Lambda's async destination, not the Scheduler DLQ. The double-retry problem: Lambda async invocations have a default internal retry count of 2 (configured in EventInvokeConfig) — combined with Scheduler's 185 default retries, a single throttle produces up to 6 Lambda invocations (3 from Lambda's internal retries × 2 Scheduler retry cycles). Disable Lambda's internal retries withaws lambda put-function-event-invoke-config --maximum-retry-attempts 0 --maximum-event-age-in-seconds 300for Scheduler-driven probes. For idempotency: usescheduleArn + scheduledTimeas the composite key (not scheduledTime alone — multiple schedules can fire at the same UTC second); implement a DynamoDB conditional put withattribute_not_exists(pk)before writing probe results. For sync invocation: use the SDK target ARNarn:aws:scheduler:::aws-sdk:lambda:invokeFunctionwithInvocationType: RequestResponse— Scheduler waits for the return value and retries on function errors, but has a 15-minute timeout hard limit imposed by Lambda. - Groups, quotas, and observability: Schedule groups provide namespace, access control, and cost tagging — create a named group per environment (
mcp-monitoring-prod,mcp-monitoring-staging) and add tags forEnvironment,Team,CostCenter, andService. The hard quota is 10,000 schedules per group; the account-region ceiling is 1,000,000 (unadjustable). For deployments approaching 10,000 schedules per environment, distribute tenants across multiple named groups — but note 500 groups per account-region is the per-group quota ceiling, and 500 × 10,000 = 5M theoretical capacity is capped by the 1M account ceiling. Cascade-delete risk:delete-schedule-groupdeletes ALL schedules in the group atomically and irreversibly — addlifecycle { prevent_destroy = true }in Terraform or CloudFormation deletion protection before any production group exists. Completed one-time schedules count toward the 10,000/group limit — run a monthly cleanup:list-schedules --state COMPLETEDthen batch-delete. Scope IAM to specific groups using the ARN resource patternarn:aws:scheduler:region:account:schedule/GROUP-NAME/*. For observability: the critical alarm isInvocationDroppedCount > 0withTreatMissingData: notBreaching— NOTmissingorbreaching, which would alarm on normal quiet periods between probe intervals. UseInvocationAttemptCountwith anomaly detection to catch accidentally disabled or deleted schedules (fires when the metric drops to zero unexpectedly — requires ~2 weeks training period). DLQ message body containsscheduleArn,scheduledTime,attemptCount,errorCode, anderrorMessage— the most common error codes areTargetLambdaPermissionError(missinglambda:InvokeFunctionin execution role),TargetThrottled(Lambda rate limit — add FlexibleTimeWindow or increase reserved concurrency), andSchedulerRolePermissionError(wrong trust principal or deleted role). Enable CloudTrail management events before you need incident investigation —DeleteScheduleGroupdoes NOT emit individualDeleteScheduleevents for the cascade, so CloudTrail is the only audit record of which group was deleted and by whom. Implement a maintenance window using SSM Parameter Store: store/mcp-monitoring/maintenance-modeas a flag, read it at Lambda startup with a 60-second TTL cache, and no-op if set — schedules keep firing (maintains metric continuity) while probes skip execution, avoiding the need to callupdate-schedulewith state DISABLED for hundreds of schedules during an outage window.
Pattern 1 — Schedule Configuration and the Trust Principal Gap
The scheduler.amazonaws.com principal and why it matters
EventBridge has three distinct invocation mechanisms that are easy to conflate: EventBridge Rules (event pattern matching on an event bus, using the events.amazonaws.com service principal), EventBridge Pipes (streaming source transformations for SQS/Kinesis/DynamoDB Streams, also on the event bus), and EventBridge Scheduler (time-based scheduled invocations, completely independent of the event bus). The Scheduler service uses its own IAM service principal: scheduler.amazonaws.com. Using the wrong principal — the one copied from a EventBridge Rules setup — is the most common first-deployment failure mode because the schedule creates without error, the IAM role exists and has the right action permissions, but the Scheduler service cannot assume the role because the trust policy excludes it. The invocation attempts happen on schedule, hit the role assumption failure, and are silently dropped. The only observable symptom is that InvocationAttemptCount stays at zero — but if you haven't pre-configured that alarm, you won't notice until your probe SLA misses.
The correct execution role trust policy for EventBridge Scheduler:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": { "Service": "scheduler.amazonaws.com" },
"Action": "sts:AssumeRole",
"Condition": {
"StringEquals": {
"aws:SourceAccount": "123456789012"
},
"ArnLike": {
"aws:SourceArn": "arn:aws:scheduler:us-east-1:123456789012:schedule/*/*"
}
}
}]
}
The SourceArn wildcard schedule/*/* covers all groups and schedule names in the account-region — this is the recommended scope for a shared execution role. The SourceAccount condition prevents confused deputy attacks where another account's Scheduler could assume this role. For per-group role isolation, narrow SourceArn to schedule/mcp-monitoring-prod/*.
Rate, cron, and one-time expressions
EventBridge Scheduler supports three expression types. Rate expressions use rate(value unit) where unit is minute, minutes, hour, hours, day, or days — singular for value 1, plural otherwise (rate(1 minute), rate(5 minutes)). Rate expressions start immediately on schedule creation and fire continuously — they're the simplest and most predictable form for regular MCP health probes.
Cron expressions use six fields: cron(minutes hours day-of-month month day-of-week year). Three critical differences from standard Unix cron: (1) there is no seconds field — the scheduler fires at the start of the specified minute; (2) the year field is required; (3) either day-of-month or day-of-week must be ? (but not both, and not neither) — specifying both as * produces a ValidationException. Example for 6 AM UTC daily: cron(0 6 * * ? *). For the first Monday of each month: cron(0 9 ? * 2#1 *).
One-time expressions use at(yyyy-mm-ddThh:mm:ss) — always in UTC, no timezone suffix. One-time schedules auto-complete after firing and transition to COMPLETED state. Always set ActionAfterCompletion: DELETE when creating one-time schedules: completed schedules count toward the 10,000/group quota and do not self-delete by default. A deployment that creates thousands of maintenance window probes without this flag will exhaust the group quota within months.
FlexibleTimeWindow and retry policy
FlexibleTimeWindow controls invocation timing within the scheduled window. With FlexibleTimeWindowMode: OFF (default), all schedules with the same interval fire simultaneously — at minute 0, all rate(5 minute) schedules across the group fire at exactly the same second, causing Lambda cold starts and function concurrency spikes. With FlexibleTimeWindowMode: FLEXIBLE and MaximumWindowInMinutes: 5, the Scheduler distributes invocations randomly within the 5-minute window. Enable this for any group with more than 50 concurrent schedules firing on the same interval.
Retry policy has two parameters: MaximumRetryAttempts (0–185, default 185) and MaximumEventAgeInSeconds (60–86400, default 86400). For MCP health probes, set MaximumEventAgeInSeconds: 300 — an invocation that failed to fire more than 5 minutes ago represents stale probe data, and retrying it at minute 7 produces a false "ok" result for a window that was actually unchecked. For ECS RunTask targets, set MaximumRetryAttempts: 0 — retrying a RunTask starts a second container instance, and the first container (which may have succeeded or partially written) remains running in parallel.
The DLQ captures invocation-level failures after retry exhaustion. The execution role needs sqs:SendMessage on the DLQ queue ARN, and the SQS queue resource policy must allow the Scheduler service principal to write:
{
"Effect": "Allow",
"Principal": { "Service": "scheduler.amazonaws.com" },
"Action": "sqs:SendMessage",
"Resource": "arn:aws:sqs:us-east-1:123456789012:mcp-scheduler-dlq",
"Condition": {
"ArnEquals": { "aws:SourceArn": "arn:aws:scheduler:us-east-1:123456789012:schedule/*/*" }
}
}
Missing the SQS queue resource policy is a silent failure: the Scheduler cannot write to the DLQ, so dropped invocations are lost with no observable record. The only way to detect missing DLQ permissions is to notice that InvocationDroppedCount is positive while InvocationsSentToDeadLetterCount stays at zero.
The update-schedule full-replacement trap
Unlike most AWS API update operations, update-schedule is a full replacement: any field not included in the call is reset to its service default. If you call update-schedule specifying only a new ScheduleExpression, the call succeeds and the expression changes — but MaximumRetryAttempts resets to 185, MaximumEventAgeInSeconds resets to 86400, DeadLetterConfig is cleared, and FlexibleTimeWindow reverts to OFF. Always retrieve the current schedule with get-schedule before updating and include all fields in the update call. Infrastructure-as-code tools that manage schedules declaratively (Terraform, CDK) handle this automatically — the risk is in manual CLI operations during incident response.
Pattern 2 — Target Integration Patterns: SDK Targets and Lambda
SDK targets and the direct API invocation model
EventBridge Scheduler can invoke over 270 AWS SDK APIs directly without a Lambda intermediary — the invocation passes a JSON input directly to the service API, which eliminates cold start latency and Lambda cost for simple scheduled operations. The target ARN format is arn:aws:scheduler:::aws-sdk:<service>:<camelCaseAction> — note the triple colon (no region or account in the ARN, because the target is the global SDK operation, not a specific resource). The input must be a JSON string matching the SDK API's request parameter structure exactly.
Context variables let you inject invocation metadata into the target input. These are template strings that the Scheduler evaluates at invocation time:
<$.context.scheduledTime>— ISO 8601 UTC timestamp of the nominal scheduled firing time (not actual invocation time, not retry time)<$.context.scheduleArn>— full ARN of the schedule<$.context.scheduleName>— schedule name only<$.context.scheduleGroupName>— group name<$.context.attemptNumber>— zero-based retry count (0 on first attempt)
Context variables use angle-bracket syntax (<$.context.scheduledTime>), not the EventBridge input transformer dollar-sign syntax. Incorrect syntax causes the variable to be passed as a literal string — the schedule creates and fires, but the downstream function receives the raw template text instead of the timestamp value.
ECS RunTask: scheduling container-based probes
The ECS RunTask SDK target (arn:aws:scheduler:::aws-sdk:ecs:runTask) is well-suited for MCP endpoint health probes where the probe logic is packaged as a Fargate container — 30-60 second tasks that probe an MCP endpoint, write results to DynamoDB, and exit. The execution role requires three IAM permissions that must all be present: ecs:RunTask on the task definition ARN, iam:PassRole on both the ECS task execution role ARN and the task role ARN (missing either PassRole triggers an AccessDenied that superficially resembles an ECS permission error), and sqs:SendMessage on the DLQ queue ARN. Pass <$.context.scheduledTime> as a container environment variable for idempotency — the task can write a DynamoDB record keyed on the scheduled time to detect and skip duplicate runs.
Set MaximumRetryAttempts: 0 for ECS RunTask targets. When Scheduler retries a RunTask that previously returned an error, it starts a second container — the first container may still be running, partially writing results. Two concurrent probe containers writing to the same DynamoDB partition key produces interleaved results. The DLQ captures the failed invocation attempt for manual investigation; the retry cost is borne by running duplicate containers in production.
Step Functions StartExecution: idempotent workflow scheduling
For MCP server workflows that involve multiple steps — probe, analyze, alert, record — Step Functions StartExecution as a Scheduler target is cleaner than a Lambda that calls Step Functions. The execution name must be unique within 90 days; use <$.context.scheduledTime> to construct it: "hourly-mcp-probe-<$.context.scheduledTime>". If Scheduler retries and the execution name already exists, Step Functions returns ExecutionAlreadyExists — this is non-fatal from the workflow perspective, but Scheduler treats it as a target error and may continue retrying.
Use Express Workflows for sub-hourly schedules: Express Workflows support up to 5-minute duration, cost per invocation (not per state transition), and are suitable for high-frequency short-lived probe workflows. Standard Workflows are appropriate for hourly or less frequent probes where duration can exceed 5 minutes or where the execution history audit trail is required.
SQS and DynamoDB SDK targets
The SQS SendMessage target (arn:aws:scheduler:::aws-sdk:sqs:sendMessage) queues a probe trigger for downstream processing — the execution role needs sqs:SendMessage on the queue ARN and sqs:GetQueueAttributes for FIFO queues. For FIFO queues, the input must include both MessageGroupId and MessageDeduplicationId — use <$.context.scheduledTime> as the deduplication ID (5-minute deduplication window matches the typical probe interval). Note that the Scheduler's DLQ and the SQS queue's own DLQ are distinct: the Scheduler DLQ captures invocation-level failures (Scheduler couldn't send the message), while the queue's DLQ captures processing failures after the message is received.
The DynamoDB PutItem target (arn:aws:scheduler:::aws-sdk:dynamodb:putItem) requires DynamoDB's typed JSON format — {"S":"value"} for strings, {"N":"42"} for numbers, not plain JSON values. Passing a plain string causes a SerializationException at invocation time. Add a ConditionExpression: "attribute_not_exists(sk)" to the input for at-most-once writes — a ConditionalCheckFailedException on retry indicates the record already exists, which is the correct behavior. Without this condition, a Scheduler retry overwrites the original probe record with stale data.
Lambda targets and the double-retry problem
Lambda targets work differently depending on invocation type. Async invocation (default InvocationType: Event): Scheduler fires the invocation and receives HTTP 202 (accepted) from Lambda — this is immediately counted as a successful invocation from Scheduler's perspective, regardless of whether the function executes successfully. Function errors go to Lambda's async destination (EventInvokeConfig.DestinationConfig.OnFailure), not the Scheduler DLQ. This asymmetry is the source of the double-retry problem.
Lambda async invocations have a built-in internal retry configured in EventInvokeConfig: the default is 2 additional attempts after the first failure, with exponential backoff over up to 6 hours. Combined with Scheduler's default of 185 retries, a single throttle produces a compounding retry cascade: Scheduler fires the invocation (attempt 1), Lambda accepts it (HTTP 202), Lambda's internal retries add 2 more attempts (attempts 2 and 3), Lambda sends to Lambda's async destination after exhaustion. Meanwhile, because Scheduler received HTTP 202 it considers the invocation successful and does not retry. The net result is 3 Lambda invocations per scheduled firing — not 6 — but if the probe is hitting a throttled endpoint, those 3 invocations all probe within a 6-hour window, poisoning the probe result timeline.
For health probes, disable Lambda's internal retry entirely:
aws lambda put-function-event-invoke-config \
--function-name mcp-health-probe \
--maximum-retry-attempts 0 \
--maximum-event-age-in-seconds 300
This means exactly one Lambda invocation per Scheduler firing. If the function throws, the invocation is lost (no internal retry). The Scheduler DLQ is unaffected — it only receives messages for invocation-level failures (Lambda throttle, wrong principal, timeout before 202 acceptance), not function execution failures. Configure Lambda's async destination for function-level failures separately.
Idempotency with scheduledTime as composite key
The <$.context.scheduledTime> variable provides the nominal scheduled firing time — this value is the same across all retry attempts for a given slot, making it suitable as an idempotency key. A conditional DynamoDB write guards against duplicate processing:
import boto3, os, json
from botocore.exceptions import ClientError
dynamodb = boto3.resource('dynamodb')
table = dynamodb.Table(os.environ['RESULTS_TABLE'])
def handler(event, context):
scheduled_time = event.get('scheduledTime')
schedule_arn = event.get('scheduleArn')
# Composite key: scheduleArn + scheduledTime
pk = f"{schedule_arn}#{scheduled_time}"
try:
table.put_item(
Item={'pk': pk, 'scheduled_time': scheduled_time, 'status': 'processing'},
ConditionExpression='attribute_not_exists(pk)'
)
except ClientError as e:
if e.response['Error']['Code'] == 'ConditionalCheckFailedException':
return {'statusCode': 200, 'body': 'duplicate, skipped'}
raise
# ... run probe ...
Use the composite key scheduleArn + scheduledTime rather than scheduledTime alone — multiple schedules may fire at the same UTC second in a large multi-tenant deployment, and scheduledTime alone would prevent the second schedule's invocation from writing its result. Log attemptNumber for debugging, but do not include it in the idempotency key.
Cross-account Lambda targets
When the Scheduler is in an operations account and the Lambda probe function is in a target account, two IAM configurations are required: a Lambda resource policy in the target account granting the scheduler.amazonaws.com service principal lambda:InvokeFunction with conditions on aws:SourceAccount and aws:SourceArn scoped to the ops account's schedule ARN pattern; and an execution role in the ops account with lambda:InvokeFunction on the target account's Lambda ARN. The execution role trust policy uses scheduler.amazonaws.com as normal. A common mistake is specifying the execution role ARN (not the service principal) in the Lambda resource policy — the Lambda service validates the calling principal as scheduler.amazonaws.com, not the assumed-role ARN.
Pattern 3 — Groups, Quotas, and Observability
Schedule group architecture for multi-tenant MCP monitoring
The default schedule group (default) is created automatically, cannot be deleted or renamed, and accepts schedules when no --group-name is specified. For production MCP monitoring, create named groups per environment: mcp-monitoring-prod, mcp-monitoring-staging, mcp-monitoring-dev. Named groups enable per-environment IAM scoping, cost allocation via tags, and group-level quota visibility.
Tag groups at creation for cost allocation — tags cannot be set after creation via create-schedule-group in some SDK versions; add them immediately with tag-resource:
aws scheduler create-schedule-group \
--name mcp-monitoring-prod \
--region us-east-1
aws scheduler tag-resource \
--resource-arn arn:aws:scheduler:us-east-1:123456789012:schedule-group/mcp-monitoring-prod \
--tags Environment=prod,Team=platform,CostCenter=eng-infra,Service=mcp-uptime
Cascade-delete and IaC protection
delete-schedule-group deletes all schedules in the group atomically and irreversibly. There is no --dry-run flag, no confirmation prompt, and no recovery path. CloudTrail records the DeleteScheduleGroup event but does NOT emit individual DeleteSchedule events for each deleted schedule — after a cascade-delete, you cannot reconstruct the deleted schedule list from CloudTrail alone.
Protect production groups in infrastructure-as-code before any schedules are created in them:
# Terraform
resource "aws_scheduler_schedule_group" "prod" {
name = "mcp-monitoring-prod"
lifecycle {
prevent_destroy = true
}
}
For CDK, use a CloudFormation deletion policy on the CfnScheduleGroup L1 construct. If a production group is accidentally deleted during an IaC teardown, you must recreate the group, recreate all schedules (from IaC state or runbooks), and accept a monitoring gap equal to the recreation time.
Quota hierarchy and scaling limits
| Quota | Default Limit | Adjustable | Notes |
|---|---|---|---|
| Schedules per account-region | 1,000,000 | No | Hard ceiling; cannot be raised via Service Quotas |
| Schedules per group | 10,000 | No | Applies to default group and every named group |
| Schedule groups per account-region | 500 | Yes | Request increase via Service Quotas |
| CreateSchedule / UpdateSchedule / DeleteSchedule TPS | 10 | Yes | ListSchedules / GetSchedule: 100 TPS |
At 500 groups × 10,000 schedules = 5,000,000 theoretical capacity, but the account ceiling of 1,000,000 is the binding constraint. For large multi-tenant deployments exceeding 200,000 schedules, distribute across multiple AWS accounts rather than increasing group count — the account ceiling is unadjustable.
One-time schedule cleanup prevents group quota exhaustion. Completed schedules persist as COMPLETED state and count toward the 10,000/group limit. List and delete them monthly:
aws scheduler list-schedules \
--group-name mcp-monitoring-prod \
--state COMPLETED \
--query 'Schedules[].Name' \
--output text | tr '\t' '\n' | while read name; do
aws scheduler delete-schedule --name "$name" --group-name mcp-monitoring-prod
done
Per-group IAM scoping
Scope IAM actions to specific groups using the resource ARN pattern:
{
"Effect": "Allow",
"Action": [
"scheduler:CreateSchedule",
"scheduler:UpdateSchedule",
"scheduler:DeleteSchedule",
"scheduler:GetSchedule",
"scheduler:ListSchedules"
],
"Resource": "arn:aws:scheduler:us-east-1:123456789012:schedule/mcp-monitoring-prod/*"
}
Without group scoping, a developer role with scheduler:* can modify schedules in any group including production. The schedule ARN format is arn:aws:scheduler:region:account:schedule/GROUP-NAME/SCHEDULE-NAME — the group name is part of the ARN, enabling resource-level permissions without tag-based conditions.
CloudWatch metrics for scheduler observability
EventBridge Scheduler publishes metrics in the AWS/Scheduler namespace with up to 5-minute delay. Six metrics are available:
| Metric | What It Counts | Alarm Strategy |
|---|---|---|
| InvocationAttemptCount | Total invocation attempts (first + retries) | Anomaly detection: alert if drops to 0 unexpectedly |
| InvocationDroppedCount | Dropped after retry exhaustion or event age exceeded | CRITICAL: > 0, TreatMissingData: notBreaching |
| InvocationThrottleCount | Scheduler-level throttling (API rate limit) | Alert if persistent (rate limit hit) |
| TargetErrorCount | Target returned error (sync targets only) | Alert for sync targets; always 0 for async Lambda |
| TargetErrorThrottledCount | Target throttle errors | Alert if persistent (target capacity problem) |
| InvocationsSentToDeadLetterCount | Messages successfully sent to DLQ | Alert on > 0 for direct DLQ monitoring |
The InvocationDroppedCount alarm is the highest-priority operational signal — a non-zero value means probe invocations are being permanently lost. The alarm configuration requires special care with TreatMissingData:
aws cloudwatch put-metric-alarm \
--alarm-name "mcp-scheduler-dropped-invocations" \
--namespace AWS/Scheduler \
--metric-name InvocationDroppedCount \
--dimensions Name=ScheduleGroup,Value=mcp-monitoring-prod \
--statistic Sum \
--period 300 \
--threshold 0 \
--comparison-operator GreaterThanThreshold \
--evaluation-periods 1 \
--treat-missing-data notBreaching \
--alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-ops-alerts
Use TreatMissingData: notBreaching. Using missing holds the alarm in its previous state (ambiguous during quiet periods). Using breaching alarms whenever no probe has fired in the evaluation period — which is expected behavior for schedules that fire every 5 minutes within a 1-minute evaluation window. The metric only appears when an invocation fires, so gaps between probe intervals produce missing data points; notBreaching treats those gaps as normal.
Anomaly detection for silenced schedules
A schedule that was accidentally disabled or deleted shows zero InvocationAttemptCount — but a fixed-threshold alarm on InvocationAttemptCount = 0 fires during normal quiet periods. Use CloudWatch anomaly detection instead:
aws cloudwatch put-metric-alarm \
--alarm-name "mcp-scheduler-silent-probe-anomaly" \
--namespace AWS/Scheduler \
--metric-name InvocationAttemptCount \
--dimensions Name=ScheduleGroup,Value=mcp-monitoring-prod \
--comparison-operator LessThanLowerOrGreaterThanUpperThreshold \
--evaluation-periods 3 \
--datapoints-to-alarm 3 \
--treat-missing-data notBreaching \
--metrics '[{"Id":"m1","MetricStat":{"Metric":{"Namespace":"AWS/Scheduler","MetricName":"InvocationAttemptCount","Dimensions":[{"Name":"ScheduleGroup","Value":"mcp-monitoring-prod"}]},"Period":300,"Stat":"Sum"}},{"Id":"e1","Expression":"ANOMALY_DETECTION_BAND(m1, 2)","Label":"InvocationAttemptCount (expected)"}]' \
--alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-ops-alerts
Anomaly detection requires approximately 2 weeks of training data before it models the baseline accurately — during the training period, suppress the alarm or accept false positives. After training, it reliably catches schedules that fire at 0 when the trained baseline predicts a positive count.
DLQ error codes and resolution patterns
Each DLQ message contains a JSON body with scheduleArn, scheduledTime, attemptCount, errorCode, errorMessage, and payload. The most common error codes:
TargetLambdaPermissionError— the execution role is missinglambda:InvokeFunctionon the target function ARN, or the Lambda resource policy is missing thescheduler.amazonaws.comprincipal. Resolution: add the missing IAM permission or update the Lambda resource policy.TargetThrottled— Lambda or the SDK target returned a throttling response. Resolution: increase Lambda reserved concurrency, add FlexibleTimeWindow to spread load, or reduce the schedule frequency.SchedulerRolePermissionError— the Scheduler cannot assume the execution role. Most commonly caused by the wrong trust principal (events.amazonaws.cominstead ofscheduler.amazonaws.com), a deleted role, or anaws:SourceArncondition that no longer matches the schedule's ARN. Resolution: verify the trust policy principal and condition values.TargetFunctionError— sync Lambda invocation returned an error. Only appears for sync targets (SDK target withInvocationType: RequestResponse) — async Lambda function errors never appear in the Scheduler DLQ. Resolution: check the Lambda function's CloudWatch Logs for the function-level error.
CloudTrail audit and the cascade-delete gap
Enable CloudTrail management events in the account-region before any production Scheduler usage — there is no retroactive capture. Scheduler API calls appear under service name scheduler.amazonaws.com. Key events to monitor: CreateSchedule, UpdateSchedule, DeleteSchedule, CreateScheduleGroup, DeleteScheduleGroup. The cascade-delete gap: when DeleteScheduleGroup fires, CloudTrail records only the group deletion event — it does not emit individual DeleteSchedule events for the hundreds or thousands of schedules deleted in the cascade. After a cascade-delete incident, the only record is the single DeleteScheduleGroup event with the deleting principal's identity; the deleted schedule names are not recoverable from CloudTrail. Maintain IaC state as the canonical source of schedule definitions, not CloudTrail.
Maintenance window pattern with SSM Parameter Store
During infrastructure maintenance, disabling hundreds of Scheduler schedules individually via update-schedule --state DISABLED is slow (10 TPS API limit), noisy in CloudTrail, and risky (the re-enable step may partially fail, leaving some schedules permanently disabled). A cleaner approach uses SSM Parameter Store as a maintenance flag:
import boto3, os, time
ssm = boto3.client('ssm')
_maintenance_cache = {'value': False, 'expires': 0}
def is_maintenance_mode():
now = time.monotonic()
if now < _maintenance_cache['expires']:
return _maintenance_cache['value']
try:
resp = ssm.get_parameter(Name='/mcp-monitoring/maintenance-mode')
val = resp['Parameter']['Value'].lower() == 'true'
except ssm.exceptions.ParameterNotFound:
val = False
_maintenance_cache.update({'value': val, 'expires': now + 60})
return val
def handler(event, context):
if is_maintenance_mode():
return {'statusCode': 200, 'body': 'maintenance mode, skipping probe'}
# ... run probe ...
Schedules keep firing during maintenance — InvocationAttemptCount continues to increment, maintaining metric continuity and preventing anomaly detection alarms from triggering on a sudden drop to zero. Probes no-op immediately after the SSM check. To start maintenance: aws ssm put-parameter --name /mcp-monitoring/maintenance-mode --value true --overwrite. To end it: aws ssm put-parameter --name /mcp-monitoring/maintenance-mode --value false --overwrite. The 60-second TTL cache means probes transition out of maintenance mode within one minute of the parameter update — no update-schedule calls required.
Consolidated failure modes
| # | Failure | Symptom | Root cause | Fix |
|---|---|---|---|---|
| 1 | Schedules create but never fire | InvocationAttemptCount = 0 | Execution role trust policy uses events.amazonaws.com not scheduler.amazonaws.com |
Update trust policy principal |
| 2 | ValidationException on cron creation | API error at create-schedule | Both day-of-month and day-of-week are * — one must be ? |
Set ? in exactly one of the two fields |
| 3 | Group quota exhaustion despite room | CreateSchedule fails with quota error | COMPLETED one-time schedules not cleaned up; count toward 10,000/group limit | Delete COMPLETED schedules; set ActionAfterCompletion: DELETE |
| 4 | Partial update resets retry policy | MaximumRetryAttempts reverts to 185 after update | update-schedule is full replacement; omitted fields reset to defaults | Always include all fields in update-schedule calls |
| 5 | ECS duplicate container instances | Two probe containers running concurrently | MaximumRetryAttempts > 0 on RunTask target; Scheduler retries start second container | Set MaximumRetryAttempts: 0 for ECS RunTask targets |
| 6 | SDK target ValidationException | API error at schedule creation | SDK action not supported, or camelCase capitalization wrong in ARN | Verify ARN against supported targets list; check capitalization |
| 7 | Context variables passed as literal strings | Function receives <$.context.scheduledTime> as text |
Incorrect syntax (using ${scheduledTime} instead of <$.context.scheduledTime>) |
Use angle-bracket syntax: <$.context.scheduledTime> |
| 8 | Lambda double-retry storm | 3+ Lambda invocations per scheduled firing | Lambda EventInvokeConfig MaximumRetryAttempts default = 2 | Set maximum-retry-attempts 0 via put-function-event-invoke-config |
| 9 | Async Lambda errors not in Scheduler DLQ | Scheduler DLQ empty despite function failures | Async invocation: HTTP 202 = Scheduler success; function errors go to Lambda destination | Configure Lambda EventInvokeConfig OnFailure destination separately |
| 10 | DynamoDB typed JSON error | SerializationException at invocation time | DynamoDB SDK target requires typed JSON format ({"S":"value"}) |
Use DynamoDB typed format in Input JSON |
| 11 | DLQ permission missing — dropped invocations lost | InvocationDroppedCount > 0, InvocationsSentToDeadLetterCount = 0 | SQS queue resource policy missing scheduler.amazonaws.com principal | Add SQS resource policy allowing scheduler.amazonaws.com SendMessage |
| 12 | Cascade-delete destroys production schedules | All schedules in group deleted | delete-schedule-group deletes all schedules atomically | Add lifecycle prevent_destroy in Terraform; use CloudFormation deletion protection |
| 13 | InvocationDroppedCount alarm false-fires | Alarm fires during quiet periods with no dropped invocations | TreatMissingData set to missing or breaching | Set TreatMissingData: notBreaching |
| 14 | Anomaly detection alarm fires during training period | False positive alarms in first 2 weeks | Model training period ~14 days; baseline not yet established | Suppress anomaly detection alarm for first 2 weeks after enabling |
| 15 | Step Functions ExecutionAlreadyExists causes retry loop | Scheduler retries indefinitely for a succeeded workflow | ExecutionAlreadyExists is treated as a target error; Scheduler retries the invocation | Use Express Workflows (no execution name uniqueness requirement); or set MaximumRetryAttempts: 0 |
Production checklists
Execution role and schedule creation
- Trust policy principal is
scheduler.amazonaws.com(notevents.amazonaws.com) - Trust policy includes both
aws:SourceAccountandaws:SourceArnconditions - Rate expression: unit is plural for values > 1 (
rate(5 minutes)notrate(5 minute)) - Cron expression: exactly one of day-of-month or day-of-week is
? - One-time schedules:
ActionAfterCompletion: DELETEset - FlexibleTimeWindow enabled for groups with >50 concurrent same-interval schedules
MaximumEventAgeInSeconds: 300for health probes- DLQ configured; execution role has
sqs:SendMessageon DLQ; DLQ has resource policy for scheduler.amazonaws.com
Target integration
- ECS RunTask:
MaximumRetryAttempts: 0; execution role hasecs:RunTask+iam:PassRoleon both task execution role and task role - Lambda async:
put-function-event-invoke-config --maximum-retry-attempts 0on probe functions - Lambda destinations (OnFailure) configured for function-level error capture
- Idempotency key is composite
scheduleArn + scheduledTime, not scheduledTime alone - DynamoDB SDK target input uses typed JSON format
- Context variables use angle-bracket syntax
<$.context.scheduledTime>
Groups and operations
- Production group has
lifecycle { prevent_destroy = true }in Terraform - Groups tagged with Environment, Team, CostCenter, Service
- Monthly COMPLETED schedule cleanup job scheduled
- IAM scoped to
schedule/GROUP-NAME/*ARN pattern per group - CloudTrail management events enabled in account-region
InvocationDroppedCount > 0alarm withTreatMissingData: notBreaching- Anomaly detection on
InvocationAttemptCount(suppressed for first 14 days) - DLQ alarm on
ApproximateNumberOfMessagesVisible > 0 - Maintenance window pattern uses SSM Parameter Store flag, not schedule state changes
Further reading
- EventBridge Scheduler setup — rate, cron, and one-time expressions, retry policy, and DLQ configuration
- SDK targets for MCP probes — ECS RunTask, Step Functions, SQS, and DynamoDB without Lambda intermediary
- Lambda targets and the double-retry problem — async vs sync, EventInvokeConfig, idempotency
- Schedule groups and quota management — cascade-delete protection, IaC patterns, multi-tenant isolation
- Observability for EventBridge Scheduler — CloudWatch metrics, DLQ alarms, CloudTrail, anomaly detection
- AWS CI/CD for MCP Servers: CodePipeline V2, CodeDeploy Blue-Green, and CDK Pipelines
- How to monitor an MCP server — health check patterns and uptime tracking
- MCP server health check — endpoint probing, timeout handling, retry strategies
- MCP server uptime monitoring — alert thresholds, SLO tracking, incident response