Guide · EventBridge Scheduler · Lambda Targets

EventBridge Scheduler Lambda Targets — Async Invocation, Idempotency, Retry, DLQ

Lambda is the most common EventBridge Scheduler target for MCP monitoring: it handles probing logic, result processing, and alerting in a single function invoked on your chosen schedule. The critical Lambda-Scheduler interaction that surprises most developers is the double-retry problem: Lambda async invocations have their own built-in retry mechanism (default: 2 additional attempts), and Scheduler has its own retry policy (default: 185 attempts over 24 hours). When both fire, a single failed schedule slot can produce up to 6 Lambda invocations — the initial attempt, 2 Lambda internal retries, and then the Scheduler retry re-starting the cycle. For MCP health probes, this means a network blip at 14:00:00 could flood your monitoring system with probe results through 14:00:06 instead of the clean single result you expected. The solution is to decide who owns retry — Scheduler or Lambda — and disable the other. The second key decision is invocation type: async (Event) for fire-and-forget probes where latency is acceptable; synchronous (RequestResponse) for at-most-once semantics where you need to know whether the invocation succeeded before Scheduler records a success.

TL;DR

Use InvocationType: Event (async) for periodic MCP health probes. Set Lambda's MaximumRetryAttempts: 0 on the async event source to disable Lambda's internal retry. Set Scheduler's MaximumRetryAttempts: 1 and MaximumEventAgeInSeconds: 300 so failures are retried once within the probe interval, then expire. Configure an SQS DLQ on the schedule. Use <$.context.scheduledTime> in the Lambda input as an idempotency key. See the core scheduler guide for execution role setup.

Invocation type: async vs synchronous

EventBridge Scheduler supports two Lambda invocation types, configured in the Input target block. The choice affects how Scheduler interprets success, how retries work, and what happens when Lambda is throttled.

InvocationTypeScheduler success conditionRetry on Lambda errorUse case
Event (async)Lambda accepts the invocation (HTTP 202)Scheduler does NOT retry on Lambda function errors — success is acceptance, not completionHealth probes, background jobs, fire-and-forget notifications
RequestResponse (sync)Lambda returns HTTP 200 with no function errorScheduler retries if Lambda returns an error response or HTTP 4xx/5xxShort-duration tasks (<15min) where you need delivery confirmation; synchronous report generation
# Async invocation (Event) — Scheduler considers it done when Lambda accepts
aws scheduler create-schedule \
  --name mcp-health-probe-5m \
  --group-name mcp-monitoring \
  --schedule-expression "rate(5 minutes)" \
  --flexible-time-window '{"Mode":"FLEXIBLE","MaximumWindowInMinutes":1}' \
  --target '{
    "Arn": "arn:aws:lambda:us-east-1:123456789012:function:mcp-health-probe",
    "RoleArn": "arn:aws:iam::123456789012:role/SchedulerExecutionRole",
    "Input": "{\"scheduledTime\":\"<$.context.scheduledTime>\",\"scheduleArn\":\"<$.context.scheduleArn>\"}",
    "RetryPolicy": {
      "MaximumRetryAttempts": 1,
      "MaximumEventAgeInSeconds": 300
    },
    "DeadLetterConfig": {
      "Arn": "arn:aws:sqs:us-east-1:123456789012:scheduler-dlq"
    }
  }'

# Note: there is no InvocationType field in the Scheduler target block itself.
# Scheduler invokes Lambda async (Event) by default when targeting a Lambda ARN.
# To use synchronous invocation, use the SDK target:
#   "Arn": "arn:aws:scheduler:::aws-sdk:lambda:invokeFunction"
# and include "InvocationType": "RequestResponse" in the Input JSON.

When using the Lambda ARN directly as the target (not the SDK target ARN), Scheduler always invokes asynchronously. To invoke synchronously, use the SDK target arn:aws:scheduler:::aws-sdk:lambda:invokeFunction and include "InvocationType": "RequestResponse" in the Input. Synchronous invocations hold the Scheduler invocation thread open until the function returns (up to 15 minutes) — Scheduler only marks the invocation as successful when the function returns with a non-error response. This is the correct mode for generating a scheduled report that must complete before the schedule is considered done.

The double-retry problem

Lambda async invocations have a built-in retry mechanism separate from Scheduler's retry policy. When a Lambda function errors (not when it is throttled or unreachable — when the function itself returns an error), Lambda internally retries the invocation up to 2 additional times before giving up. This internal retry system is configured on the Lambda function's EventInvokeConfig, not on the Scheduler schedule.

# Lambda's internal retry for async invocations (EventInvokeConfig)
# Default: MaximumRetryAttempts=2, MaximumEventAgeInSeconds=21600 (6 hours)

# Disable Lambda's internal retry — let Scheduler own all retry logic
aws lambda put-function-event-invoke-config \
  --function-name mcp-health-probe \
  --maximum-retry-attempts 0 \
  --maximum-event-age-in-seconds 300

# Alternatively, disable Scheduler retry — let Lambda own all retry logic
# (Set RetryPolicy.MaximumRetryAttempts=0 on the schedule)

# What happens with both retry systems active (default):
# 14:00:00 — Scheduler invokes Lambda (attempt 1)
# 14:00:00 — Lambda function errors
# 14:00:01 — Lambda retries internally (attempt 2)
# 14:00:01 — Lambda function errors again
# 14:00:03 — Lambda retries internally (attempt 3, exponential backoff)
# 14:00:03 — Lambda function errors again → Lambda sends to Lambda DLQ
# → Scheduler receives HTTP 202 (async acceptance) — considers invocation SUCCESSFUL
# → Scheduler does NOT retry because async acceptance = success
# Result: Lambda DLQ has 1 message; Scheduler DLQ has 0; CloudWatch shows 0 failures

The subtle issue with async invocations: Scheduler records a success as soon as Lambda accepts the invocation (HTTP 202), regardless of whether the function succeeded. Lambda's internal retry fires after that acceptance — Scheduler never sees the function error. This means Scheduler's retry policy is effectively irrelevant for async Lambda invocations when the failure is a function-level error (exception in the handler). Scheduler retry only activates for invocation-level failures: Lambda throttling (HTTP 429), service unavailability (HTTP 503), or permissions errors.

For synchronous invocations via the SDK target, both retry systems interact differently: Scheduler waits for the function to return, sees the function error in the response, and applies its own retry policy. Lambda's internal retry does not apply to synchronous invocations. Set MaximumRetryAttempts only on Scheduler for sync invocations.

Idempotency with scheduledTime

Scheduler can invoke a Lambda function more than once for the same scheduled slot in edge cases: Scheduler retry, Lambda internal retry, and infrastructure-level at-least-once delivery guarantees. MCP health probe handlers must be idempotent — writing the same probe result twice must not corrupt the monitoring dataset.

# Pass context variables into Lambda input for idempotency
--target '{
  "Arn": "arn:aws:lambda:...",
  "Input": "{
    \"scheduledTime\": \"<$.context.scheduledTime>\",
    \"scheduleArn\": \"<$.context.scheduleArn>\",
    \"attemptNumber\": \"<$.context.attemptNumber>\"
  }"
}'

# Lambda handler using scheduledTime as idempotency key
import boto3, json, hashlib

dynamodb = boto3.resource('dynamodb')
table = dynamodb.Table('mcp-probe-results')

def handler(event, context):
    scheduled_time = event['scheduledTime']      # e.g. "2026-10-04T14:00:00Z"
    schedule_arn = event['scheduleArn']
    attempt = int(event.get('attemptNumber', '0'))

    # Idempotency key: (schedule ARN, scheduled slot time)
    # Same slot retried → same key → conditional write fails → skip duplicate
    result_id = hashlib.sha256(
        f"{schedule_arn}#{scheduled_time}".encode()
    ).hexdigest()[:16]

    # Run the actual health probe
    probe_result = probe_mcp_endpoint()

    try:
        table.put_item(
            Item={
                'pk': f'probe#{result_id}',
                'sk': scheduled_time,
                'attempt': attempt,
                'status': probe_result['status'],
                'latency_ms': probe_result['latency_ms'],
                'ttl': 1757980800  # expire after 90 days
            },
            ConditionExpression='attribute_not_exists(pk)'
        )
    except dynamodb.meta.client.exceptions.ConditionalCheckFailedException:
        # Idempotent: this slot was already processed
        print(f"Duplicate invocation for slot {scheduled_time}, skipping")
        return {'statusCode': 200, 'body': 'duplicate'}

    return {'statusCode': 200, 'body': 'ok'}

The scheduledTime variable contains the nominal scheduled time — the time the schedule was supposed to fire, not the actual invocation time. This is the correct key for idempotency: two invocations for the same scheduled slot will have the same scheduledTime value, even if Scheduler fired them 30 seconds apart due to flexible time window or retry delays. The attemptNumber can be logged for debugging but should not be part of the idempotency key — two attempts for the same slot should both be deduplicated.

DLQ configuration for Lambda targets

There are two DLQ systems that can apply to a Scheduler → Lambda invocation, and they capture different failure types:

DLQConfigured onCapturesMessage format
Scheduler DLQSchedule target DeadLetterConfigInvocation failures: Lambda throttle, service unavailability, execution role errors, MaximumRetryAttempts exhaustedScheduler envelope with scheduleArn, scheduledTime, attemptNumber, error code
Lambda destination (async)Lambda function EventInvokeConfig.DestinationConfig.OnFailureFunction-level failures after all Lambda internal retries exhausted (handler errors, timeouts, OOM)Lambda async invocation record with request/response details
# Configure both DLQ layers for full coverage
# 1. Scheduler DLQ — catches invocation-level failures
aws scheduler update-schedule \
  --name mcp-health-probe-5m \
  --group-name mcp-monitoring \
  # (all fields) ...
  --target '{
    ...
    "DeadLetterConfig": {
      "Arn": "arn:aws:sqs:us-east-1:123456789012:scheduler-dlq"
    }
  }'

# 2. Lambda async destination — catches function-level failures
aws lambda put-function-event-invoke-config \
  --function-name mcp-health-probe \
  --maximum-retry-attempts 0 \
  --destination-config '{
    "OnFailure": {
      "Destination": "arn:aws:sqs:us-east-1:123456789012:lambda-failure-dlq"
    }
  }'

# Lambda destination also supports SNS, EventBridge, Lambda as destinations
# Use SNS to fan-out to PagerDuty + Slack on Lambda function failure

For async invocations, the Scheduler DLQ will be empty on function-level errors (handler exceptions) because Scheduler sees HTTP 202 acceptance as success. Only Lambda's destination captures function-level failures for async invocations. For synchronous invocations, only the Scheduler DLQ is relevant — Lambda's async destination does not apply to synchronous calls.

Cross-account Lambda targets

EventBridge Scheduler can invoke Lambda functions in a different AWS account. This is useful for centralized monitoring: a Scheduler schedule in an operations account invokes health probe functions deployed in each application account.

# Cross-account Lambda target setup

# 1. In the target account (application account): add Lambda resource policy
#    allowing the scheduler execution role from the ops account to invoke
aws lambda add-permission \
  --function-name mcp-health-probe \
  --statement-id AllowSchedulerFromOpsAccount \
  --action lambda:InvokeFunction \
  --principal scheduler.amazonaws.com \
  --source-account "111111111111" \
  --source-arn "arn:aws:scheduler:us-east-1:111111111111:schedule/mcp-monitoring/*"

# 2. In the ops account: create execution role with cross-account InvokeFunction
{
  "Statement": [{
    "Effect": "Allow",
    "Action": "lambda:InvokeFunction",
    "Resource": "arn:aws:lambda:us-east-1:222222222222:function:mcp-health-probe"
  }]
}

# 3. In the ops account: create schedule with cross-account Lambda ARN
aws scheduler create-schedule \
  --name mcp-probe-app-account \
  --group-name mcp-monitoring \
  --schedule-expression "rate(5 minutes)" \
  --flexible-time-window '{"Mode":"FLEXIBLE","MaximumWindowInMinutes":1}' \
  --target '{
    "Arn": "arn:aws:lambda:us-east-1:222222222222:function:mcp-health-probe",
    "RoleArn": "arn:aws:iam::111111111111:role/SchedulerExecutionRole",
    ...
  }'

The Lambda resource policy in the target account must grant permissions to scheduler.amazonaws.com (not to the execution role ARN directly). Use --source-account and --source-arn to scope which Scheduler schedules are authorized — without these conditions, any Scheduler schedule in any account that knows your Lambda ARN could invoke it.

Cross-account DLQ is more complex: the Scheduler DLQ must be in the same account as the schedule (the ops account), but the Lambda destination for function-level failures is in the target account. Plan your alerting accordingly — you need visibility into both queues to cover all failure modes.

Lambda version and alias targeting

# Target a specific Lambda version (immutable)
"Arn": "arn:aws:lambda:us-east-1:123456789012:function:mcp-health-probe:42"

# Target an alias (mutable — update alias to change which version fires)
"Arn": "arn:aws:lambda:us-east-1:123456789012:function:mcp-health-probe:production"

# The execution role must allow InvokeFunction on both the function and the alias/version:
{
  "Effect": "Allow",
  "Action": "lambda:InvokeFunction",
  "Resource": [
    "arn:aws:lambda:us-east-1:123456789012:function:mcp-health-probe",
    "arn:aws:lambda:us-east-1:123456789012:function:mcp-health-probe:*"
  ]
}
# The :* wildcard covers all versions and aliases

Using an alias (:production) is the recommended pattern for production schedules — you can update which Lambda version the alias points to without modifying the schedule. Updating the schedule to point to a new function ARN requires an update-schedule call with all fields, which risks accidentally changing other schedule parameters. Alias promotion is a single aws lambda update-alias call scoped only to the function.

Failure modes reference

FailureSymptomFix
Async invocation: function errors not in Scheduler DLQLambda DLQ has messages; Scheduler DLQ empty; Scheduler shows successExpected: async acceptance = Scheduler success; Lambda function errors go to Lambda's async destination, not Scheduler DLQ — configure Lambda destination separately
Double-retry storm (6× invocations per slot)DynamoDB has duplicate probe records; unexpected Lambda concurrency spikesBoth Lambda internal retry and Scheduler retry active; disable one: set Lambda MaximumRetryAttempts=0 or Scheduler MaximumRetryAttempts=0
Synchronous invocation timeoutScheduler marks invocation failed; retries; Lambda function still runningSynchronous invocations over 15 minutes will always fail (Lambda hard limit); use async target or redesign for shorter duration; add FlexibleTimeWindow to reduce concurrent retry pile-up
Cross-account resource policy wrong principalLambda returns 403 despite correct execution role permissionsLambda resource policy must grant scheduler.amazonaws.com service principal, not the execution role ARN directly — resource policy and IAM policy are both required for cross-account
Idempotency key collisionProbe results overwritten for same slot across multiple schedulesInclude scheduleArn in the idempotency key — scheduledTime alone is not unique if multiple schedules fire at the same time; use scheduleArn + scheduledTime as composite key