Guide · AWS EventBridge Scheduler · Scheduled Automation

AWS EventBridge Scheduler for MCP Servers — Scheduled Health Checks, Rate Expressions, Execution Role

AWS EventBridge Scheduler is a fully managed scheduler that invokes targets on a rate, cron, or one-time schedule — completely separate from EventBridge Rules. For MCP server operators, Scheduler solves the "run a health probe every 5 minutes" problem without managing CloudWatch Events, Lambda triggers, or cron on EC2. Scheduler directly invokes over 270 AWS SDK API actions as first-class targets, meaning you can run an ECS task, start a Step Functions execution, or write a DynamoDB record on a schedule with no intermediary Lambda. The critical architectural distinction: EventBridge Rules react to events (pattern matching against an event bus); EventBridge Scheduler drives scheduled time-based invocations independently of any event bus. They are different services with different APIs, IAM namespaces, and quota systems. Scheduler's execution role must explicitly trust scheduler.amazonaws.com — the EventBridge Rules principal (events.amazonaws.com) is not accepted. Key decisions: choosing between rate, cron, and one-time expression types; configuring FlexibleTimeWindow to spread load; setting retry policy and DLQ to handle target failures; and picking the right target type for your use case.

TL;DR

Create a rate-expression schedule: aws scheduler create-schedule --name mcp-health-check-5m --schedule-expression "rate(5 minutes)" --flexible-time-window '{"Mode":"FLEXIBLE","MaximumWindowInMinutes":1}' --target '{"Arn":"arn:aws:lambda:us-east-1:123456789012:function:mcp-health-probe","RoleArn":"arn:aws:iam::123456789012:role/SchedulerExecutionRole","RetryPolicy":{"MaximumRetryAttempts":1,"MaximumEventAgeInSeconds":300},"DeadLetterConfig":{"Arn":"arn:aws:sqs:us-east-1:123456789012:scheduler-dlq"}}'. The execution role trust policy must name scheduler.amazonaws.com, not events.amazonaws.com. See the Lambda target guide for idempotency patterns and the SDK target guide for calling AWS APIs directly.

EventBridge Scheduler vs EventBridge Rules vs EventBridge Pipes

Three EventBridge services are commonly confused. They solve different problems and are not interchangeable:

ServiceTriggerUse caseIAM principal
EventBridge RulesEvent pattern match on event busReact to AWS service events (S3 upload, EC2 state change, CodePipeline stage change)events.amazonaws.com
EventBridge SchedulerTime-based schedule (rate, cron, one-time)Periodic MCP health probes, scheduled reports, key rotation triggers, maintenance windowsscheduler.amazonaws.com
EventBridge PipesStreaming/queue source (SQS, Kinesis, DynamoDB Streams)Filter and enrich records flowing between a queue/stream and a target; event transformation pipelinepipes.amazonaws.com

Mixing up the IAM principal is the most common setup error. If you create an execution role with a trust policy that names events.amazonaws.com and attach it to a Scheduler schedule, the schedule creation may succeed but invocations will fail with an access denied error — Scheduler cannot assume a role not in its trust relationship.

The second common confusion is thinking EventBridge Scheduler replaces CloudWatch Events. It does — AWS migrated scheduled CloudWatch Events (old cron-expression rules on the default event bus) to EventBridge Scheduler. Existing CloudWatch Events scheduled rules still work but new scheduled automation should use the Scheduler API, which has better quotas, flexible time windows, DLQ support, and direct SDK target integration.

Schedule expression types

EventBridge Scheduler supports three expression types. Choose based on whether the invocation is recurring or one-time, and whether the exact minute matters:

TypeSyntaxExampleUse case
raterate(value unit)rate(5 minutes)Fixed-interval recurring work — health probes, metric collection
croncron(min hr dom mon dow yr)cron(0 6 * * ? *)Calendar-aligned work — daily reports at 6 AM UTC, weekly cleanups on Sunday
one-time (at)at(yyyy-mm-ddThh:mm:ss)at(2026-12-01T02:00:00)Scheduled maintenance window, future rotation trigger
# rate — every 5 minutes
--schedule-expression "rate(5 minutes)"

# cron — every day at 06:00 UTC (field order: min hr dom mon dow yr)
--schedule-expression "cron(0 6 * * ? *)"
# Note: EventBridge Scheduler cron uses a 6-field format (no seconds field)
# The '?' in the day-of-week field means "no specific value" — required when dom is '*'

# cron — every Monday at 08:00 UTC
--schedule-expression "cron(0 8 ? * MON *)"

# one-time — fire once at a specific UTC time
--schedule-expression "at(2026-12-01T02:00:00)"
# Timezone-naive: always UTC. Set --schedule-expression-timezone for local time.

# one-time with timezone (IANA timezone)
--schedule-expression "at(2026-12-01T02:00:00)"
--schedule-expression-timezone "America/New_York"

The cron field order for Scheduler differs from standard Unix cron: it is minute hour day-of-month month day-of-week year — there is no seconds field, and the year field is required (use * for every year). The ? character means "no specific value" and must appear in either the day-of-month or day-of-week field — not both, and not neither. Omitting it causes a ValidationException: Expression is not valid error that does not explain which field is wrong.

One-time schedules fire once and then enter the COMPLETED state. They are not deleted automatically unless you set --action-after-completion DELETE. Accumulated completed one-time schedules count against your quota (10,000 schedules per group). For high-volume one-time scheduling, always set ActionAfterCompletion: DELETE.

Execution role

Every schedule requires an execution role — the IAM role that Scheduler assumes to invoke your target. The trust policy must grant scheduler.amazonaws.com the ability to assume the role. The role permissions must allow the specific action on the target resource.

# Trust policy — must name scheduler.amazonaws.com, NOT events.amazonaws.com
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Principal": {
        "Service": "scheduler.amazonaws.com"
      },
      "Action": "sts:AssumeRole",
      "Condition": {
        "StringEquals": {
          "aws:SourceAccount": "123456789012"
        },
        "ArnLike": {
          "aws:SourceArn": "arn:aws:scheduler:us-east-1:123456789012:schedule/*/*"
        }
      }
    }
  ]
}

# Permission policy for Lambda target
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "lambda:InvokeFunction",
      "Resource": [
        "arn:aws:lambda:us-east-1:123456789012:function:mcp-health-probe",
        "arn:aws:lambda:us-east-1:123456789012:function:mcp-health-probe:*"
      ]
    },
    {
      "Effect": "Allow",
      "Action": "sqs:SendMessage",
      "Resource": "arn:aws:sqs:us-east-1:123456789012:scheduler-dlq"
    }
  ]
}

The aws:SourceArn condition in the trust policy is optional but strongly recommended — it prevents confused deputy attacks where a schedule in another account could trick your role into invoking targets in your account. The wildcard pattern arn:aws:scheduler:us-east-1:123456789012:schedule/*/* covers schedules in any group; narrow it to a specific group if you maintain per-environment roles.

The DLQ permission (sqs:SendMessage on the DLQ ARN) must be in the same execution role as the target permission. If Scheduler cannot write to the DLQ after a failed invocation, the event is silently dropped — you lose both the invocation and the failure record. Always include the DLQ send permission when you configure a DLQ.

FlexibleTimeWindow

FlexibleTimeWindow tells Scheduler it may invoke the target any time within a window around the scheduled time, rather than at the exact moment. This spreads load across many concurrent schedules and reduces the risk of Lambda cold start pile-ups when hundreds of MCP health checks are scheduled at the same interval.

# OFF — invoke exactly at schedule time (default)
--flexible-time-window '{"Mode":"OFF"}'

# FLEXIBLE — invoke within 0–5 minutes after scheduled time
--flexible-time-window '{"Mode":"FLEXIBLE","MaximumWindowInMinutes":5}'

# Example: health check scheduled at 12:00, window=5 → fires between 12:00 and 12:05
# Scheduler chooses the actual time; you cannot predict or control it within the window

# For rate(5 minutes) with window=1:
# - Schedule fires "at or up to 1 minute after" each 5-minute interval
# - Effective interval is still ~5 minutes but jittered ±1 minute

Use FLEXIBLE for health probes and batch jobs where exact timing is not critical. Keep the window small (1–5 minutes) for near-real-time work. Use OFF for time-sensitive operations like scheduled maintenance windows where you need the action at a specific moment, or for scheduled reports where stakeholders expect the 06:00 UTC delivery time.

FlexibleTimeWindow does not apply to one-time (at) schedules — they always fire at the specified time.

Retry policy and DLQ

Scheduler's retry policy controls what happens when the target invocation fails — for example, when Lambda returns an error, when the SDK target API returns a throttle, or when the execution role cannot assume its permissions temporarily. Retry policy settings are per-schedule, not global.

# Full retry + DLQ configuration in the target block
--target '{
  "Arn": "arn:aws:lambda:us-east-1:123456789012:function:mcp-health-probe",
  "RoleArn": "arn:aws:iam::123456789012:role/SchedulerExecutionRole",
  "RetryPolicy": {
    "MaximumRetryAttempts": 2,
    "MaximumEventAgeInSeconds": 3600
  },
  "DeadLetterConfig": {
    "Arn": "arn:aws:sqs:us-east-1:123456789012:scheduler-dlq"
  }
}'

# MaximumRetryAttempts: 0–185 (default: 185)
# MaximumEventAgeInSeconds: 60–86400 seconds (default: 86400 = 24 hours)
# Scheduler retries with exponential backoff until either limit is hit

# For health probes: set MaximumEventAgeInSeconds=300 (5 minutes)
# This ensures a failed probe attempt doesn't retry into the next probe window
FieldDefaultRecommended for health probesWhy
MaximumRetryAttempts1851–2Probes should fail fast and let the next scheduled probe retry; 185 retries over 24 hours creates false alert suppression
MaximumEventAgeInSeconds86400300A stale probe result from 20 minutes ago is useless; expire it before the next scheduled window
DLQNoneRequiredWithout DLQ, dropped invocations are invisible — only CloudWatch metrics reveal failures, and only if you have alarms

The DLQ receives a message for each invocation that exhausts retries or exceeds the event age limit. The message body contains the original input payload plus Scheduler metadata (scheduleArn, scheduledTime, attempt number). Monitoring the DLQ depth via CloudWatch is the simplest operational alarm for Scheduler failures — see the observability guide for alarm patterns.

One gotcha: Scheduler's retry policy is independent of Lambda's built-in async invocation retry (Lambda retries async invocations up to 2 times by default). If you invoke Lambda asynchronously with Scheduler, both retry systems activate: Lambda retries internally, and if Lambda ultimately fails, Scheduler's retry kicks in and re-invokes Lambda. Set Lambda's MaximumRetryAttempts: 0 on the async event source configuration if you want only Scheduler's retry logic to apply — otherwise you get up to 3× Lambda retries × 2 Scheduler retries = 6 total invocations per scheduled slot.

Creating and managing schedules

# Create a schedule
aws scheduler create-schedule \
  --name mcp-health-check-5m \
  --group-name mcp-monitoring \
  --schedule-expression "rate(5 minutes)" \
  --flexible-time-window '{"Mode":"FLEXIBLE","MaximumWindowInMinutes":1}' \
  --target '{
    "Arn": "arn:aws:lambda:us-east-1:123456789012:function:mcp-health-probe",
    "RoleArn": "arn:aws:iam::123456789012:role/SchedulerExecutionRole",
    "Input": "{\"endpoint\":\"https://api.example.com/mcp\",\"timeout\":10}",
    "RetryPolicy": {
      "MaximumRetryAttempts": 1,
      "MaximumEventAgeInSeconds": 300
    },
    "DeadLetterConfig": {
      "Arn": "arn:aws:sqs:us-east-1:123456789012:scheduler-dlq"
    }
  }' \
  --state ENABLED

# Disable a schedule (pause without deleting)
aws scheduler update-schedule \
  --name mcp-health-check-5m \
  --group-name mcp-monitoring \
  --state DISABLED \
  # ... (all other fields required on update)

# IMPORTANT: update-schedule requires ALL fields, not just the changed one.
# Use get-schedule to retrieve current config, then modify and re-submit.

# Get schedule details (includes last-run information)
aws scheduler get-schedule \
  --name mcp-health-check-5m \
  --group-name mcp-monitoring

# List schedules in a group
aws scheduler list-schedules \
  --group-name mcp-monitoring \
  --query 'Schedules[*].{Name:Name,State:State,Expr:ScheduleExpression}'

# Delete a schedule
aws scheduler delete-schedule \
  --name mcp-health-check-5m \
  --group-name mcp-monitoring

The --state parameter accepts ENABLED, DISABLED. Use DISABLED during maintenance windows to temporarily pause probes without deleting the schedule configuration. The update-schedule API requires all parameters — there is no partial-update PATCH semantics. Retrieve the current schedule with get-schedule first, modify the fields you want to change, and submit the full config.

Failure modes reference

FailureSymptomFix
Execution role trust principal wrongSchedule creates successfully but invocations never happen; no CloudWatch metric activityTrust policy names events.amazonaws.com instead of scheduler.amazonaws.com — Scheduler cannot assume the role; update trust policy to use correct principal
Cron expression ValidationExceptionValidationException: Expression is not valid on create-scheduleMissing or misplaced ? in cron fields — either day-of-month or day-of-week must be ?; 6-field format (no seconds); year field required
DLQ silent drop on failureFailed invocations disappear; no DLQ messagesExecution role missing sqs:SendMessage on DLQ ARN — Scheduler cannot write failure records; add SQS permission to execution role
Lambda double-retry stormLambda invoked 6× per schedule slot instead of 2×Lambda async retry (default 2 attempts) compounds Scheduler retry; set Lambda event source MaximumRetryAttempts: 0 or set Scheduler MaximumRetryAttempts: 0
One-time schedule quota exhaustionConflictException: Schedule quota exceededCompleted one-time schedules not deleted; set ActionAfterCompletion: DELETE or periodically delete COMPLETED schedules with list-schedules + delete-schedule
update-schedule overwrites configSchedule changes unexpectedly (target, expression reset)update-schedule is a full replacement — always get-schedule first, modify fields, resubmit all fields; partial updates silently lose unspecified fields