Guide · AWS Lambda · Lambda Layers

Lambda Layer Versioning, Immutability, and Update Strategies for MCP Servers

Every publish-layer-version call creates a new, permanently immutable version — the existing version is never overwritten, the ZIP it contains cannot change, and referencing layer-arn:5 always gives you the exact same bytes it did the day you published it. This immutability is the foundation of safe dependency management for MCP servers: you can update a layer and test it on a canary function while production functions keep pinning the previous version, then promote by updating the function configuration. For building the layer artifact itself see shared dependencies guide; for sharing layers across accounts see cross-account layers guide.

TL;DR

Layer versions are immutable integers starting at 1. Functions reference a specific version ARN — there is no "latest" pointer. When you publish a new version, existing functions are unaffected until you explicitly update their configuration. The safe update workflow is: publish new version → update one canary function → invoke and smoke-test → batch-update remaining functions if green. Never delete a layer version that any function references. Use RemovalPolicy.RETAIN in CDK and RetentionPolicy: Retain in SAM to prevent accidental deletion. Keep at least 3 previous versions available for quick rollback.

Versioning model

Lambda layer versions are auto-incremented integers starting at 1. Each publish-layer-version call increments the counter by 1 regardless of how small the change was. The version number is part of the ARN: arn:aws:lambda:us-east-1:123456789012:layer:mcp-shared-deps:7. There is no concept of a "latest" ARN or a mutable alias that points to the current version — every reference is a specific, pinned version number.

This design is intentional and important for MCP servers: it means that updating a layer to fix a security vulnerability in zod cannot break a production function unless you explicitly update that function's layer ARN. The functions you have not updated keep running on the previous layer version indefinitely. This is the opposite of how many package managers work, where a ^3.0.0 range can pull in a breaking patch release on the next npm ci. Lambda layers are always deterministic.

# List all published versions of a layer
aws lambda list-layer-versions \
  --layer-name mcp-shared-deps \
  --region us-east-1 \
  --query 'LayerVersions[*].{Version:Version,ARN:LayerVersionArn,Created:CreatedDate,Description:Description}' \
  --output table
# ------------------------------------------------------------------------------------------------
# |                               ListLayerVersions                                              |
# +---------+-------------------------------+------------------+----------------------------------+
# | Version | ARN                           | Created          | Description                      |
# +---------+-------------------------------+------------------+----------------------------------+
# | 8       | arn:...:mcp-shared-deps:8     | 2026-10-10T09:00 | v8: zod 3.23.8, jose 5.9.6       |
# | 7       | arn:...:mcp-shared-deps:7     | 2026-09-15T14:22 | v7: aws-sdk 3.650.0 bump         |
# | 6       | arn:...:mcp-shared-deps:6     | 2026-08-01T11:05 | v6: added @aws-sdk/lib-dynamodb  |
# | 5       | arn:...:mcp-shared-deps:5     | 2026-07-12T08:44 | v5: security patch jose 5.9.0    |
# | 4       | arn:...:mcp-shared-deps:4     | 2026-06-03T16:30 | v4: initial stable set           |
# +---------+-------------------------------+------------------+----------------------------------+

# Get detailed metadata for a specific layer version
aws lambda get-layer-version \
  --layer-name mcp-shared-deps \
  --version-number 8 \
  --region us-east-1
# {
#   "LayerVersionArn": "arn:aws:lambda:us-east-1:123456789012:layer:mcp-shared-deps:8",
#   "Version": 8,
#   "Description": "v8: zod 3.23.8, jose 5.9.6",
#   "CreatedDate": "2026-10-10T09:00:00.000+0000",
#   "CompatibleRuntimes": ["nodejs20.x", "nodejs22.x"],
#   "CompatibleArchitectures": ["x86_64", "arm64"],
#   "Content": {
#     "Location": "https://prod-04-2014-tasks.s3.us-east-1.amazonaws.com/...",  # pre-signed URL
#     "CodeSha256": "abc123...",
#     "CodeSize": 19234567,
#     "UnzippedCodeSize": 84123456
#   }
# }

The version ARN is permanent. Even if you delete a layer version, the ARN cannot be reused — version 8 is version 8 forever, even after deletion. This means you can use the ARN as a stable identifier in audit logs, deployment records, and rollback scripts. A function's CloudWatch Logs will show which layer version ARN was active at the time of each invocation if you emit it as a structured log field.

# Confirm which layer version a function is currently using
aws lambda get-function-configuration \
  --function-name mcp-tool-s3 \
  --region us-east-1 \
  --query 'Layers'
# [
#   {
#     "Arn": "arn:aws:lambda:us-east-1:123456789012:layer:mcp-shared-deps:7",
#     "CodeSize": 19100000
#   }
# ]
# This function is still on version 7 — it has NOT been updated to version 8 yet

# Check all MCP tool functions to see which layer version each uses
aws lambda list-functions \
  --region us-east-1 \
  --query 'Functions[?starts_with(FunctionName, `mcp-tool-`)].{Name:FunctionName,Layers:Layers}' \
  --output json | jq '.[] | {name: .Name, layers: (.Layers // [] | map(.Arn))}'

Version pinning strategy

The recommended pinning strategy for MCP server deployments is to keep production functions pinned to a known-good layer version and only update them after a canary function running the new version has been validated. Staging and development functions can always track the latest published version — their role is precisely to discover breakage before it reaches production.

Pinning is not optional in production — it is automatic. Lambda never updates a function's layer reference unless you explicitly call update-function-configuration or deploy via CDK/SAM. A new layer version published to your account has zero effect on running functions until you update their configuration. This is different from, say, npm packages where a loose version range in package.json can pull in a new version on the next npm ci.

# Check which layer version each MCP function is pinned to
# Useful before a layer update to know the blast radius
MCP_FUNCTIONS=(
  mcp-tool-s3
  mcp-tool-dynamodb
  mcp-tool-secrets
  mcp-tool-sqs
  mcp-tool-sns
  mcp-tool-ec2
  mcp-tool-iam
  mcp-tool-cloudwatch
  mcp-tool-bedrock
  mcp-tool-sts
)

echo "Function name | Layer ARN"
echo "------------- | ---------"
for fn in "${MCP_FUNCTIONS[@]}"; do
  LAYER=$(aws lambda get-function-configuration \
    --function-name "$fn" \
    --region us-east-1 \
    --query 'Layers[0].Arn' \
    --output text 2>/dev/null || echo "no layer")
  echo "$fn | $LAYER"
done
# Promote all functions to a new layer version after canary validation
# Run this ONLY after the canary smoke test has passed
NEW_LAYER_ARN="arn:aws:lambda:us-east-1:123456789012:layer:mcp-shared-deps:8"

for fn in "${MCP_FUNCTIONS[@]}"; do
  echo "Updating $fn to $NEW_LAYER_ARN..."
  aws lambda update-function-configuration \
    --function-name "$fn" \
    --layers "$NEW_LAYER_ARN" \
    --region us-east-1 \
    --no-cli-pager

  # Wait for the configuration update to propagate before moving to the next function
  # update-function-configuration is eventually consistent; invoking before it
  # completes may still use the old configuration
  aws lambda wait function-updated \
    --function-name "$fn" \
    --region us-east-1

  echo "$fn updated successfully"
done
echo "All functions updated to layer version 8."

If your MCP server has more than ~10 functions, the sequential batch update above may be slow. You can parallelize the update-function-configuration calls, but Lambda rate-limits concurrent configuration updates per account. The safe parallelism is roughly 5–10 concurrent updates; beyond that you will hit TooManyRequestsException. For large fleets, use exponential backoff or consider managing layer updates through CDK which handles the concurrency limits automatically.

Safe update workflow

A layer update in a multi-function MCP server has the potential for a wide blast radius: a single bad dependency bump that gets promoted to all 10 functions simultaneously takes down every tool at once. The canary workflow prevents this by exposing the new layer version to exactly one function first, validating it, and then promoting all remaining functions only after the canary is green.

# Step 1: Build and publish the new layer version
cd layer-build/nodejs
npm install zod@3.23.8 jose@5.9.6 @aws-sdk/client-s3@3.700.0
cd ..
zip -r layer-v8.zip nodejs/

NEW_LAYER_ARN=$(aws lambda publish-layer-version \
  --layer-name mcp-shared-deps \
  --description "v8: zod 3.23.8, jose 5.9.6, AWS SDK 3.700.0" \
  --zip-file fileb://layer-v8.zip \
  --compatible-runtimes nodejs20.x nodejs22.x \
  --compatible-architectures x86_64 arm64 \
  --region us-east-1 \
  --query 'LayerVersionArn' \
  --output text)

echo "New layer ARN: $NEW_LAYER_ARN"
# arn:aws:lambda:us-east-1:123456789012:layer:mcp-shared-deps:8
# Step 2: Update the canary function only
# Choose the lowest-traffic MCP tool as the canary — mcp-tool-ec2 if it gets few calls
CANARY_FUNCTION="mcp-tool-ec2"

aws lambda update-function-configuration \
  --function-name "$CANARY_FUNCTION" \
  --layers "$NEW_LAYER_ARN" \
  --region us-east-1 \
  --no-cli-pager

# Wait for the update to complete — IMPORTANT: do not invoke until this returns
aws lambda wait function-updated \
  --function-name "$CANARY_FUNCTION" \
  --region us-east-1

echo "Canary $CANARY_FUNCTION updated to layer version 8. Running smoke test..."
# Step 3: Invoke the canary function with a test payload and check the response
# Use a minimal valid MCP tools/call invocation
TEST_PAYLOAD='{"jsonrpc":"2.0","method":"tools/call","id":1,"params":{"name":"ec2-describe-instances","arguments":{"region":"us-east-1","maxResults":1}}}'

RESPONSE=$(aws lambda invoke \
  --function-name "$CANARY_FUNCTION" \
  --payload "$TEST_PAYLOAD" \
  --region us-east-1 \
  /tmp/canary-response.json \
  --query 'StatusCode' \
  --output text)

echo "HTTP status: $RESPONSE"
cat /tmp/canary-response.json | jq .

# Check for Lambda runtime errors (HTTP 200 but error in body)
if cat /tmp/canary-response.json | jq -e '.errorMessage' > /dev/null 2>&1; then
  echo "CANARY FAILED: Lambda runtime error in response body"
  echo "Do NOT proceed with batch update. Investigate the error."
  cat /tmp/canary-response.json | jq '.errorMessage, .errorType, .stackTrace'
  exit 1
fi

# Check for MCP-level error
if cat /tmp/canary-response.json | jq -e '.isError == true' > /dev/null 2>&1; then
  echo "CANARY WARNING: MCP tool returned isError=true — check if this is expected for the test input"
  cat /tmp/canary-response.json | jq '.content'
fi

echo "Canary smoke test passed. Proceeding with batch update."
# Step 4: Batch-update remaining functions (canary already updated — skip it)
REMAINING_FUNCTIONS=(
  mcp-tool-s3
  mcp-tool-dynamodb
  mcp-tool-secrets
  mcp-tool-sqs
  mcp-tool-sns
  mcp-tool-iam
  mcp-tool-cloudwatch
  mcp-tool-bedrock
  mcp-tool-sts
)

for fn in "${REMAINING_FUNCTIONS[@]}"; do
  echo "Updating $fn..."
  aws lambda update-function-configuration \
    --function-name "$fn" \
    --layers "$NEW_LAYER_ARN" \
    --region us-east-1 \
    --no-cli-pager
  aws lambda wait function-updated --function-name "$fn" --region us-east-1
  echo "$fn -> layer version 8"
done

echo "All functions updated. Layer v8 is now in production."
# Step 5 (if canary failed): No action needed — canary still uses old ARN
# Production functions were never touched. To confirm:
for fn in "${MCP_FUNCTIONS[@]}"; do
  LAYER=$(aws lambda get-function-configuration \
    --function-name "$fn" \
    --region us-east-1 \
    --query 'Layers[0].Arn' \
    --output text)
  echo "$fn: $LAYER"
done
# mcp-tool-ec2: arn:...:layer:mcp-shared-deps:8   (canary — affected)
# mcp-tool-s3:  arn:...:layer:mcp-shared-deps:7   (production — unchanged)
# mcp-tool-dynamodb: arn:...:layer:mcp-shared-deps:7  (production — unchanged)
# ... all others still on version 7

CDK-native layer versioning

CDK's LayerVersion construct uses a SHA-256 content hash of the asset directory to determine whether a new layer version should be published. If the asset content has not changed since the last deploy, CDK reuses the existing layer version and does not call publish-layer-version. This makes layer updates incremental and efficient — rebuilding the CDK stack after changing only a Lambda function handler will not republish the layer.

CDK also tracks the layer version ARN attachment per function. When a new layer version is published (because the asset content hash changed), CDK automatically updates all functions in the stack that reference that layer construct to point at the new version ARN. This eliminates the manual batch-update step — but it also means all functions are updated in a single cdk deploy, which is the opposite of the canary approach above. To get canary behavior in CDK, separate the canary function into a different CDK stack or use Lambda aliases (see the canary deployments section below).

// CDK: LayerVersion construct with RETAIN policy and explicit version tracking
import * as lambda from "aws-cdk-lib/aws-lambda";
import * as cdk from "aws-cdk-lib";
import * as path from "path";
import { Construct } from "constructs";

export class McpLayerStack extends cdk.Stack {
  // Expose the layer so other stacks can reference it
  public readonly sharedDepsLayer: lambda.LayerVersion;

  constructor(scope: Construct, id: string, props?: cdk.StackProps) {
    super(scope, id, props);

    this.sharedDepsLayer = new lambda.LayerVersion(this, "McpSharedDeps", {
      layerVersionName: "mcp-shared-deps",
      code: lambda.Code.fromAsset(path.join(__dirname, "../layers/shared-deps")),
      compatibleRuntimes: [
        lambda.Runtime.NODEJS_20_X,
        lambda.Runtime.NODEJS_22_X,
      ],
      compatibleArchitectures: [
        lambda.Architecture.X86_64,
        lambda.Architecture.ARM_64,
      ],
      description: "Shared deps: AWS SDK v3, zod, jose",
      removalPolicy: cdk.RemovalPolicy.RETAIN,
      // RETAIN: if this construct is removed from the stack, or if the CDK deploy
      // tries to delete the old version during an update, it is kept in Lambda.
      // This prevents functions that still reference the old ARN from breaking.
    });

    // Export the layer ARN as a CloudFormation output for cross-stack references
    new cdk.CfnOutput(this, "LayerArn", {
      value: this.sharedDepsLayer.layerVersionArn,
      exportName: "McpSharedDepsLayerArn",
      description: "ARN of the latest deployed mcp-shared-deps layer version",
    });
  }
}
// CDK: Referencing the layer in the functions stack
// Using a cross-stack reference so the layer can be deployed independently
import * as lambda from "aws-cdk-lib/aws-lambda";
import * as cdk from "aws-cdk-lib";
import { Construct } from "constructs";
import { McpLayerStack } from "./layer-stack";

export class McpFunctionsStack extends cdk.Stack {
  constructor(scope: Construct, id: string, layerStack: McpLayerStack, props?: cdk.StackProps) {
    super(scope, id, props);

    const toolNames = [
      "mcp-tool-s3",
      "mcp-tool-dynamodb",
      "mcp-tool-secrets",
    ];

    for (const toolName of toolNames) {
      new lambda.Function(this, toolName, {
        functionName: toolName,
        runtime: lambda.Runtime.NODEJS_22_X,
        handler: "index.handler",
        code: lambda.Code.fromAsset(
          path.join(__dirname, `../tools/${toolName}/dist`)
        ),
        layers: [layerStack.sharedDepsLayer],
        // CDK automatically tracks the layerVersionArn and updates the function
        // configuration when a new layer version is published by McpLayerStack
      });
    }
  }
}

// app.ts entry point
const app = new cdk.App();
const layerStack = new McpLayerStack(app, "McpLayerStack", { env: { region: "us-east-1" } });
const fnStack = new McpFunctionsStack(app, "McpFunctionsStack", layerStack, { env: { region: "us-east-1" } });
// fnStack depends on layerStack implicitly via the cross-stack reference
# Trigger a layer update: modify the dependency in layers/shared-deps/nodejs/package.json
# and run npm install, then redeploy both stacks
cd layers/shared-deps/nodejs
npm install zod@3.23.8  # update to new version
cd ../../..

# Deploy the layer stack first — this publishes the new layer version
cdk deploy McpLayerStack

# Check what CDK will change before deploying functions
cdk diff McpFunctionsStack
# Stack McpFunctionsStack
# Resources
# [~] AWS::Lambda::Function mcp-tool-s3
#   |-- [~] Layers
#   |     |-- [~] .0:
#   |           |-- [-] arn:aws:lambda:us-east-1:123456789012:layer:mcp-shared-deps:7
#   |           |-- [+] arn:aws:lambda:us-east-1:123456789012:layer:mcp-shared-deps:8

# Deploy functions stack — all functions updated to layer v8
cdk deploy McpFunctionsStack

Canary deployments with layers

Lambda function aliases and weighted routing let you direct a fraction of traffic to a function version that uses a new layer while the majority of traffic continues to a version pinned to the old layer. This is the most surgical form of canary testing for MCP tools: a real percentage of production tool calls exercises the new dependency, with automatic rollback if the error rate rises.

The key insight is that function versions (not just function configurations) are also immutable. Publishing a function version captures the function code AND the current layer ARN references at that moment. Version 3 of the function always uses layer version 4; version 4 of the function always uses layer version 5. Aliases point to function versions, and weighted routing splits traffic between aliases.

# Publish a function version pinned to layer version 4 (current production)
# First confirm the function is on layer v4
aws lambda update-function-configuration \
  --function-name mcp-tool-s3 \
  --layers "arn:aws:lambda:us-east-1:123456789012:layer:mcp-shared-deps:4" \
  --region us-east-1
aws lambda wait function-updated --function-name mcp-tool-s3 --region us-east-1

# Publish as a numbered version (immutable snapshot of code + config + layer ARN)
STABLE_VERSION=$(aws lambda publish-version \
  --function-name mcp-tool-s3 \
  --description "Stable: layer v4, code v3" \
  --region us-east-1 \
  --query 'Version' \
  --output text)
echo "Stable function version: $STABLE_VERSION"  # e.g., 3
# Update function to layer v5 and publish as the canary version
aws lambda update-function-configuration \
  --function-name mcp-tool-s3 \
  --layers "arn:aws:lambda:us-east-1:123456789012:layer:mcp-shared-deps:5" \
  --region us-east-1
aws lambda wait function-updated --function-name mcp-tool-s3 --region us-east-1

CANARY_VERSION=$(aws lambda publish-version \
  --function-name mcp-tool-s3 \
  --description "Canary: layer v5, code v3" \
  --region us-east-1 \
  --query 'Version' \
  --output text)
echo "Canary function version: $CANARY_VERSION"  # e.g., 4
# Create a production alias pointing to the stable version
aws lambda create-alias \
  --function-name mcp-tool-s3 \
  --name production \
  --function-version "$STABLE_VERSION" \
  --description "Production traffic — layer v4" \
  --region us-east-1

# Add 10% canary traffic to function version 4 (layer v5)
# 90% goes to STABLE_VERSION (layer v4); 10% goes to CANARY_VERSION (layer v5)
aws lambda update-alias \
  --function-name mcp-tool-s3 \
  --name production \
  --routing-config "AdditionalVersionWeights={\"$CANARY_VERSION\": 0.1}" \
  --region us-east-1

# Verify the routing configuration
aws lambda get-alias \
  --function-name mcp-tool-s3 \
  --name production \
  --region us-east-1
# {
#   "Name": "production",
#   "FunctionVersion": "3",
#   "RoutingConfig": {
#     "AdditionalVersionWeights": { "4": 0.1 }
#   }
# }
# After monitoring canary metrics: if green, promote to 100%
# Point the alias entirely to the canary version and remove the weight split
aws lambda update-alias \
  --function-name mcp-tool-s3 \
  --name production \
  --function-version "$CANARY_VERSION" \
  --routing-config "AdditionalVersionWeights={}" \
  --region us-east-1

# If canary shows errors: roll back by removing the weight split immediately
# (just set routing-config back to empty — all traffic returns to STABLE_VERSION)
aws lambda update-alias \
  --function-name mcp-tool-s3 \
  --name production \
  --function-version "$STABLE_VERSION" \
  --routing-config "AdditionalVersionWeights={}" \
  --region us-east-1
# Rollback is instantaneous — next invocation uses stable version with layer v4

The MCP client calling your tool must invoke the alias ARN (ending in :production) rather than the plain function name or a versioned ARN. Lambda resolves the weighted routing at invoke time. Your API Gateway integration or function URL should target the alias ARN to participate in the canary split.

Retention and lifecycle

Lambda never automatically deletes layer versions — they accumulate indefinitely. This is a feature, not a bug: it ensures that functions deployed months ago continue to work even if you have published 50 new layer versions since. However, old layer versions do consume storage quota, and in large accounts with many layers and many versions, cleaning up safely becomes a periodic maintenance task.

Before deleting any layer version, you must verify that no function references it. Lambda does not enforce a referential integrity constraint — it will allow you to delete a layer version that a function depends on, and the function will silently fail the next time it cold-starts. The fix is to always audit before deleting and to prefer retaining rather than deleting.

# Find all functions that use a specific layer version before considering deletion
# This is the critical safety check — never skip it
LAYER_VERSION_ARN="arn:aws:lambda:us-east-1:123456789012:layer:mcp-shared-deps:4"

# List all functions and filter for those that reference this layer version
aws lambda list-functions \
  --region us-east-1 \
  --query "Functions[?Layers != null && contains(Layers[*].Arn, '$LAYER_VERSION_ARN')].FunctionName" \
  --output text
# If empty: no functions reference this version — safe to delete
# If non-empty: update those functions to a newer layer version first
#!/usr/bin/env python3
# safe-layer-cleanup.py
# Lists layer versions and identifies which are safe to delete
# (not referenced by any function in the account)

import boto3
import json

REGION = "us-east-1"
LAYER_NAME = "mcp-shared-deps"
RETAIN_LAST_N = 3  # Always keep at least 3 versions for quick rollback

lambda_client = boto3.client("lambda", region_name=REGION)

# Get all versions of the layer
versions_response = lambda_client.list_layer_versions(LayerName=LAYER_NAME)
versions = versions_response["LayerVersions"]
# Sort descending by version number
versions.sort(key=lambda v: v["Version"], reverse=True)
print(f"Total versions of {LAYER_NAME}: {len(versions)}")

# Collect all layer ARNs referenced by any function in the account
referenced_arns = set()
paginator = lambda_client.get_paginator("list_functions")
for page in paginator.paginate():
    for fn in page["Functions"]:
        for layer in fn.get("Layers", []):
            referenced_arns.add(layer["Arn"])

print(f"\nLayer ARNs referenced by at least one function: {len(referenced_arns)}")

# Identify safe-to-delete candidates:
# - Not in the most recent RETAIN_LAST_N versions
# - Not referenced by any function
recent_arns = {v["LayerVersionArn"] for v in versions[:RETAIN_LAST_N]}
print(f"\nRetaining the {RETAIN_LAST_N} most recent versions unconditionally.")
print("\nCandidate versions for deletion:")
print(f"{'Version':<10} {'ARN':<65} {'Referenced':<12} {'Safe to delete'}")
print("-" * 100)

for version in versions:
    arn = version["LayerVersionArn"]
    ver_num = version["Version"]
    is_recent = arn in recent_arns
    is_referenced = arn in referenced_arns
    safe = not is_recent and not is_referenced

    status = "YES" if safe else "NO"
    ref_status = "YES" if is_referenced else "no"
    recent_status = "(recent)" if is_recent else ""
    print(f"{ver_num:<10} {arn:<65} {ref_status:<12} {status} {recent_status}")

    # Uncomment to actually delete safe versions:
    # if safe:
    #     lambda_client.delete_layer_version(LayerName=LAYER_NAME, VersionNumber=ver_num)
    #     print(f"  Deleted version {ver_num}")
# Manually delete a specific layer version (only after confirming no references)
# This is PERMANENT and cannot be undone
aws lambda delete-layer-version \
  --layer-name mcp-shared-deps \
  --version-number 2 \
  --region us-east-1
# No output on success — the version is gone

# Verify deletion
aws lambda get-layer-version \
  --layer-name mcp-shared-deps \
  --version-number 2 \
  --region us-east-1
# Error: An error occurred (ResourceNotFoundException) ...
# Confirms the version is deleted — and cannot be recovered

AliveMCP and version-correlated downtime

When a layer update causes an MCP tool to fail, the failure signal appears in AliveMCP before you see it in CloudWatch Logs. AliveMCP detects the first failing probe within 60 seconds of the layer update completing — before a human has time to check the console or before the first CloudWatch alarm fires (which typically requires 3–5 data points over 3–5 minutes). The combination of AliveMCP's timestamp and your deployment timestamp makes root cause analysis immediate: if the MCP tool went down at 14:23 and you deployed layer version 8 at 14:22, the correlation is obvious.

Configure AliveMCP with your MCP server's function URL or API Gateway endpoint, a test MCP tools/call payload, and a Slack webhook on your engineering channel. When you run a layer update, you will see the "mcp-tool-s3 is up" notification within 60 seconds confirming the update succeeded — or a "mcp-tool-s3 went down" alert if something broke. AliveMCP also monitors the HTTP status code, the response body structure, and response time — so a regression from 200ms to 4,000ms cold-start after a layer update shows up as a latency alert, not just a downtime alert.

You can use AliveMCP's status API to programmatically gate the canary promotion step: if the AliveMCP check for the canary function is green after 5 minutes, proceed with the batch update; if it is red, abort. This turns AliveMCP into an automated canary gate rather than a manual observation tool.

# Shell script: gate the canary promotion on AliveMCP status
# Assumes ALIVEMCP_API_KEY and CHECK_ID env vars are set
# Replace with your actual AliveMCP check ID for the canary function

ALIVEMCP_STATUS=$(curl -s \
  -H "Authorization: Bearer $ALIVEMCP_API_KEY" \
  "https://api.alivemcp.com/v1/checks/$CHECK_ID/status" \
  | jq -r '.status')

echo "AliveMCP canary status: $ALIVEMCP_STATUS"

if [ "$ALIVEMCP_STATUS" = "up" ]; then
  echo "Canary is healthy. Proceeding with batch update..."
  # run batch update loop
else
  echo "Canary is $ALIVEMCP_STATUS. Aborting layer promotion."
  echo "Check AliveMCP dashboard for details: https://alivemcp.com/dashboard"
  exit 1
fi

Failure modes

SymptomCauseFix
Function silently uses old layer after publishing a new versionFunction configuration still references the old layer version ARN — Lambda never auto-updates function-to-layer referencesExplicitly update function with aws lambda update-function-configuration --function-name $fn --layers $NEW_LAYER_ARN; CDK handles this automatically on cdk deploy when the layer asset hash changes
ResourceNotFoundException: Layer version does not exist on cold start or configuration updateLayer version was deleted while a function still referenced it — either via delete-layer-version or via CDK's RemovalPolicy.DESTROY during a stack updateRe-publish from the same ZIP if available (creates a new version number, then update function ARN); going forward use RemovalPolicy.RETAIN in CDK and RetentionPolicy: Retain in SAM for all layer constructs
All MCP tool functions broken simultaneously after "routine" dependency bumpShared layer was updated and all functions were updated at once without canary testing; a breaking API change in the new dependency version caused all functions to fail togetherThe canary workflow prevents this: always update one function first, smoke-test it, then batch-update the rest; pin production functions to a tested layer version; only promote after canary validation
TooManyRequestsException when batch-updating functionsAWS rate-limits concurrent update-function-configuration calls per account; updating 10 functions in rapid succession without waiting hits the limitAdd aws lambda wait function-updated between each update in the loop (already shown in the batch update examples above); alternatively use CDK which handles concurrency limits internally
CDK drift: function uses a different layer version than CDK expects after cdk diffManual aws lambda update-function-configuration --layers was run outside CDK, causing the actual state to diverge from the CDK-managed state; CDK's next diff sees the function's layer ARN as needing an update back to the CDK-tracked versionAlways use CDK for layer updates in CDK-managed stacks; to recover from drift, run cdk diff to see what CDK wants to change, then cdk deploy to reconcile; or import the current state with cdk import if you want to keep the manually-set version