AWS Lambda · 2026-10-10 · Lambda Layers arc
Lambda Layers for MCP Servers: Shared Dependencies, Version Management, Cross-Account Distribution, and Custom Runtimes
Four Lambda Layers topics composed into three structural patterns for MCP server deployments — shared dependencies via Lambda Layers (layer ZIP must have nodejs/ as the top-level directory — not node_modules/ directly — because the Node.js runtime prepends /opt/nodejs/node_modules to NODE_PATH, so import { z } from 'zod' resolves transparently at runtime; mark all layer packages external in esbuild with --external:zod --external:"@aws-sdk/*" — omitting this flag bundles deps into both the ZIP and the layer, silently negating the size reduction; per-function ZIP drops from 20–30 MB to under 1 MB and cold start from 800–1,500 ms to 300–600 ms; split into three layers by update frequency: AWS SDK layer stable monthly, validation/auth layer for security patches, shared utils layer for weekly churn; RemovalPolicy.RETAIN in CDK and RetentionPolicy: Retain in SAM are critical — without them a stack update can delete a layer version that functions still reference, causing ResourceNotFoundException at the next cold start; 50 MB compressed direct upload limit, 250 MB unzipped total across function plus all layers), Lambda layer versioning and update strategies (every publish-layer-version creates a new immutable auto-incremented integer — there is no "latest" ARN or mutable alias; functions pin to a specific version ARN and remain on it indefinitely until you explicitly call update-function-configuration; safe 5-step update workflow: publish new version → update one low-traffic canary function → invoke with a test MCP tools/call payload and check for errorMessage in the response body as well as the JSON-RPC status code → batch-update remaining functions only if canary is green → rollback is zero-cost because non-canary functions were never touched; CDK uses a SHA-256 content hash of the layer asset directory and rebuilds only on change, then auto-updates all functions in the stack in a single cdk deploy — to preserve canary behavior in CDK, separate the canary function into a different stack; function aliases with AdditionalVersionWeights split traffic at 10% canary without requiring two separate CDK stacks; before deleting any layer version run the Python audit script that collects all function layer ARNs and identifies which versions are safe to delete — Lambda does not enforce referential integrity; retain at least 3 previous versions for quick rollback), cross-account Lambda layer sharing (layer permissions are resource-based policies on specific versions — new versions start with empty policies, no inheritance from previous versions; grant lambda:GetLayerVersion to a specific account ID or use --principal "*" --principal-org-id o-xxx for org-wide access; the consumer account needs no IAM policy on the platform account's layer ARN — the access check runs against the layer's resource-based policy in the platform account only; layer permissions are strictly per-region — replicate both publish-layer-version and add-layer-version-permission to every region your MCP platform operates in using a shell loop; revocation timing: warm execution environments continue until recycled, then the next cold start fails with AccessDeniedException: lambda:GetLayerVersion; never revoke the previous layer version until confirming all consumer functions have migrated via a function configuration audit; deprecation window of 2–4 weeks communicated via Slack announcement with the new ARN and changelog), and custom runtimes via Lambda Layers (a custom runtime is a Lambda Layer containing a bootstrap executable at the ZIP root — /opt/bootstrap in the execution environment; Lambda Runtime API contract: GET /runtime/invocation/next blocks until next event, POST to /runtime/invocation/{requestId}/response returns result — the loop handles both warm and cold paths; bootstrap must be executable before zipping — chmod +x, then verify with unzip -l layer.zip | grep bootstrap; Bun 1.1.x layer: 40 MB x86_64 binary, 80–200 ms cold start, JavaScriptCore engine, native TypeScript, Bun.S3Client replaces @aws-sdk/client-s3; Deno 2.x layer: 90 MB binary typically requires S3 upload path, 150–350 ms cold start, explicit permission flags --allow-env --allow-net --allow-read=/var/task in bootstrap; Python 3.13 layer: CPython compiled in the Lambda base Docker image (public.ecr.aws/lambda/python:3.12 builder), 70 MB compressed via S3, --prefix=/opt/python313 install; architecture-specific builds are mandatory — Bun and Deno publish separate binaries for linux-x64 and linux-aarch64; function runtime set to provided.al2023, not nodejs22.x; Provisioned Concurrency at 2–5 instances eliminates cold starts for latency-critical MCP tools sitting in tight agent loops). 10-row consolidated failure modes table.
Pattern 1: The shared dependency contract — ZIP layout, runtime resolution, and bundler externals
A typical MCP server built on Lambda has 5–15 tool functions, each independently deployable but sharing a common set of heavy dependencies: the AWS SDK v3 family of clients, zod for input validation, and jose for JWT verification. Without layers, each function ZIP contains a full private copy of every package it imports — the same 20 MB of @aws-sdk and zod bytes repeated across every tool. With layers, those bytes live once at /opt/nodejs/node_modules and are shared by all functions attached to the layer. The storage and cold-start math quickly follows: 10 functions at 25 MB each becomes a 25 MB layer plus ~10 MB of handler ZIPs combined. The shared dependencies guide covers the build and attachment workflow in detail.
The ZIP structure is non-negotiable. Lambda's Node.js runtime prepends /opt/nodejs/node_modules to NODE_PATH before your handler initializes. For this automatic path injection to work, the layer ZIP must have nodejs/ as its top-level directory — nodejs/node_modules/zod/, not node_modules/zod/ directly at the ZIP root. If you zip node_modules/ directly, the runtime mounts the layer at /opt and the module resolution algorithm never finds your packages. The diagnosis is a runtime Cannot find module 'zod' error even though the layer is listed in the function configuration — the layer is attached, the bytes are present, but at the wrong path.
# Correct structure — nodejs/ is at the ZIP root
mkdir -p layer-build/nodejs
cd layer-build/nodejs
npm init -y
npm install zod@3.23.8 jose@5.9.6 \
@aws-sdk/client-s3@3.650.0 \
@aws-sdk/client-secrets-manager@3.650.0 \
@aws-sdk/client-dynamodb@3.650.0
cd ..
zip -r layer.zip nodejs/
# Verify: first entries must be nodejs/, not node_modules/
unzip -l layer.zip | head -5
# nodejs/
# nodejs/node_modules/
# nodejs/node_modules/zod/
# ✓ correct
# Incorrect: zipping from inside the nodejs/ directory
# zip -r layer.zip node_modules/
# Would produce: node_modules/zod/ -- Lambda cannot find it
Bundler externals are the other half of the contract. When you build your handler with esbuild, the default behavior is to bundle every import into the output file. If you add a layer for zod but do not mark it external in the bundler, esbuild bundles zod into the function ZIP anyway. The layer is still mounted at /opt and the function runs correctly — but the size reduction is completely negated. The correct esbuild invocation for a handler that relies on a shared layer marks every layer package as external:
# Build the handler without bundling layer packages
npx esbuild mcp-tool-s3/index.ts \
--bundle \
--platform=node \
--target=node22 \
--external:zod \
--external:"@aws-sdk/*" \
--external:jose \
--outfile=dist/index.js \
--minify
cd dist && zip -r ../function.zip index.js && cd ..
ls -lh function.zip
# ~150 KB — versus ~25 MB if --external flags were omitted
Split layers by update frequency, not by function. A single monolithic layer containing every dependency has maximum blast radius: any change to any package requires rebuilding the entire layer and updating every function that references it. A stable structure for MCP servers is three layers: the AWS SDK layer (changes monthly at most, ~80 MB unzipped for a v3 client selection), a validation and auth layer (zod, jose, ajv — updated more often for security patches but small at ~5 MB), and a shared internal utilities layer (your own compiled TypeScript, auth middleware, response formatters — changes most frequently, smallest size). The AWS SDK layer is listed first in --layers and the utilities layer is listed last so it can shadow anything from the first layer on a name conflict.
IaC retention policies are mandatory, not optional. CDK's LayerVersion construct and SAM's AWS::Serverless::LayerVersion both have a "what happens to the old layer version when we publish a new one" policy setting. The default in both tools is DESTROY. Lambda does not enforce referential integrity — it will allow you to delete a layer version that running functions still depend on. The deletion succeeds silently; the functions fail on their next cold start with ResourceNotFoundException. The safe default for any production MCP server is RemovalPolicy.RETAIN in CDK and RetentionPolicy: Retain in SAM on all layer constructs, and a manual audit before any manual deletion of old layer versions.
Pattern 2: The version lifecycle contract — immutability, canary updates, and CDK cross-stack patterns
Lambda layer versions are auto-incremented integers starting at 1. arn:aws:lambda:us-east-1:123456789012:layer:mcp-shared-deps:7 contains exactly the same bytes it did the day it was published — forever. There is no "latest" pointer, no mutable alias that dereferences to the current version, no mechanism for a layer version to receive updates after publication. This immutability is the foundation for safe dependency management in a multi-function MCP server: you can update the layer for one function while all others continue running on the previous version indefinitely. The versioning guide covers the full lifecycle in detail.
Functions never self-update their layer ARN. Publishing a new layer version has zero effect on any running function until you call update-function-configuration on that function with the new layer ARN. This is the opposite of how package managers work — there is no semver range that will pull in a new layer version on the next deploy. A function configured with layer version 7 will use layer version 7 for every invocation until you explicitly change it. This means a layer security patch published by your platform team does not protect functions that have not been updated to reference the new version. Auditing which functions are on which version before and after a security update is a production requirement.
The canary workflow prevents fleet-wide failures from bad dependency bumps. A shared layer that gets promoted to all 10 MCP tool functions simultaneously has a blast radius of 10 — a breaking API change in the new layer version takes down every tool at once. The canary workflow limits exposure to one function during validation:
# Step 1: Publish new layer version
NEW_LAYER_ARN=$(aws lambda publish-layer-version \
--layer-name mcp-shared-deps \
--description "v8: zod 3.23.8, jose 5.9.6, AWS SDK 3.700.0" \
--zip-file fileb://layer-v8.zip \
--compatible-runtimes nodejs20.x nodejs22.x \
--compatible-architectures x86_64 arm64 \
--region us-east-1 \
--query 'LayerVersionArn' --output text)
# Step 2: Update the lowest-traffic MCP tool only
aws lambda update-function-configuration \
--function-name mcp-tool-ec2 \
--layers "$NEW_LAYER_ARN" \
--region us-east-1
aws lambda wait function-updated --function-name mcp-tool-ec2 --region us-east-1
# Step 3: Smoke-test — check for Lambda runtime errors in the response BODY
# HTTP 200 does NOT mean the function worked; a Lambda error payload is also HTTP 200
RESPONSE=$(aws lambda invoke \
--function-name mcp-tool-ec2 \
--payload '{"jsonrpc":"2.0","method":"tools/call","id":1,"params":{"name":"ec2-describe-instances","arguments":{"region":"us-east-1","maxResults":1}}}' \
--region us-east-1 /tmp/canary-response.json \
--query 'StatusCode' --output text)
# Check for runtime error (errorMessage key in body = Lambda threw before returning)
if cat /tmp/canary-response.json | jq -e '.errorMessage' > /dev/null 2>&1; then
echo "FAILED: $(cat /tmp/canary-response.json | jq -r '.errorMessage')"
echo "Canary stuck on layer version 8. All other functions unchanged."
exit 1
fi
# Step 4 (if green): Batch-update remaining functions
for fn in mcp-tool-s3 mcp-tool-dynamodb mcp-tool-secrets mcp-tool-sqs; do
aws lambda update-function-configuration --function-name "$fn" --layers "$NEW_LAYER_ARN" --region us-east-1
aws lambda wait function-updated --function-name "$fn" --region us-east-1
done
The canary approach has a critical property: if the smoke test fails, no production functions were touched. The canary function is on the new layer version, all others remain on the old version. There is no rollback step for production — it was never updated. The only remediation needed is to restore the canary function to the previous layer version.
CDK's cross-stack pattern for layer updates. CDK computes a SHA-256 content hash of the layer asset directory. If the content has not changed, CDK reuses the existing layer version and skips the publish-layer-version call — incremental deploys of handlers do not republish the layer. When the content does change, CDK publishes a new layer version and auto-updates every function in the stack that references the layer construct. To preserve canary behavior in CDK without manual CLI steps, separate the canary function into its own stack so it can be deployed before the main functions stack:
// Cross-stack: layer stack exports the ARN; functions stack consumes it
export class McpLayerStack extends cdk.Stack {
public readonly sharedDepsLayer: lambda.LayerVersion;
constructor(scope: Construct, id: string, props?: cdk.StackProps) {
super(scope, id, props);
this.sharedDepsLayer = new lambda.LayerVersion(this, "McpSharedDeps", {
layerVersionName: "mcp-shared-deps",
code: lambda.Code.fromAsset(path.join(__dirname, "../layers/shared-deps")),
compatibleRuntimes: [lambda.Runtime.NODEJS_20_X, lambda.Runtime.NODEJS_22_X],
compatibleArchitectures: [lambda.Architecture.X86_64, lambda.Architecture.ARM_64],
removalPolicy: cdk.RemovalPolicy.RETAIN,
});
new cdk.CfnOutput(this, "LayerArn", {
value: this.sharedDepsLayer.layerVersionArn,
exportName: "McpSharedDepsLayerArn",
});
}
}
// Deploy order: cdk deploy McpLayerStack → cdk deploy McpCanaryStack → validate
// → cdk deploy McpFunctionsStack
Weighted alias routing enables surgical canary testing. Function aliases with AdditionalVersionWeights split live traffic between two function versions — one pinned to the old layer, one pinned to the new. A 10% canary weight means 1 in 10 real MCP tool invocations exercises the new dependency set. The alias ARN is what your API Gateway integration or Function URL targets; the routing decision happens at invoke time inside Lambda. Rollback is instantaneous: set the alias routing config back to 100% on the stable version.
Retention and lifecycle. Lambda never automatically deletes layer versions — they accumulate indefinitely. Before deleting any version, run a Python audit that collects all layer ARNs referenced by any function in the account and compares them against candidate versions. Never delete a version that any function references. The practical guidance is: retain the 3 most recent versions unconditionally for quick rollback, and then audit and clean up anything older.
Pattern 3: Cross-account distribution and custom runtimes
A Lambda layer in account A cannot be used by a function in account B unless you explicitly grant permission. The permission mechanism is a resource-based policy on the specific layer version — the same model as S3 bucket policies or Lambda function resource policies. There is no automatic inheritance: publishing version 8 of a layer does not inherit the permissions you granted on version 7. Every new version starts with an empty policy and must be granted separately. The cross-account layers guide covers the full permission model, and the custom runtimes guide covers deploying Bun, Deno, and Python 3.13 via custom runtime layers.
Organization-wide grants eliminate per-account grant management. For MCP platform teams managing many product accounts, granting each account individually creates a maintenance burden: every new product team must be added, every new layer version requires a new set of add-layer-version-permission calls. The org-wide pattern — --principal "*" --principal-org-id o-xxx — eliminates this. Any account currently in the organization at the time a function is updated to reference the layer will be permitted. New accounts added to the org later get access automatically; there is no need to re-run the permission grant when the org grows.
# Grant org-wide access on the new layer version (run in the PLATFORM account)
aws lambda add-layer-version-permission \
--layer-name mcp-shared-deps \
--version-number 8 \
--statement-id allow-entire-org \
--action lambda:GetLayerVersion \
--principal "*" \
--principal-org-id o-xxxxxxxxxxxx \
--region us-east-1
# Consumer account: attach the cross-account layer by its full ARN
# The consumer needs no IAM policy on the platform account's layer --
# access is granted by the resource-based policy on the layer itself
aws lambda update-function-configuration \
--function-name mcp-tool-s3 \
--layers arn:aws:lambda:us-east-1:111122223333:layer:mcp-shared-deps:8 \
--region us-east-1
# CDK in the consumer account: import by ARN, no cross-account IAM needed
const sharedDepsLayer = lambda.LayerVersion.fromLayerVersionArn(
this, "SharedDepsLayer",
"arn:aws:lambda:us-east-1:111122223333:layer:mcp-shared-deps:8"
);
Layer permissions are strictly per-region. A layer ARN in us-east-1 cannot be attached to a function in eu-west-1 — Lambda requires the function and all its layers to be in the same region. For MCP platforms operating across regions, the platform CI/CD pipeline must replicate both the publish-layer-version call and the add-layer-version-permission call to every region. The same ZIP can be used for the publish call in each region (pure JavaScript packages are architecture-neutral); the permission grants must target the correct region-specific ARN in each case.
Revocation timing requires coordination. When you run remove-layer-version-permission on a layer version, warm execution environments in consumer accounts continue to work — they already have the layer bytes loaded. The failure happens at the next cold start, when Lambda tries to fetch the layer content from the platform account and gets AccessDeniedException: lambda:GetLayerVersion. This means revocation causes intermittent failures that become total failures after traffic subsides and all warm environments are recycled. Before revoking any version, audit consumer functions, communicate a deprecation date, and confirm migration via a function configuration audit.
Custom runtimes: the bootstrap contract. A custom runtime layer is a Lambda Layer that contains an executable named bootstrap at the ZIP root — not in a subdirectory, not in bin/. Lambda mounts the layer at /opt, so /opt/bootstrap is the path it executes when the function runtime is set to provided.al2023. The executable must have the Unix execute bit set before zipping; the chmod +x bootstrap step inside the Docker build or shell script is required. A bootstrap that is not executable causes Runtime.InvalidEntrypoint immediately at cold start.
The Lambda Runtime API contract that every bootstrap must implement:
#!/bin/sh
# Minimal bootstrap illustrating the Lambda Runtime API contract
# Replace the python3 invocation with exec to your target runtime
RUNTIME_API="http://${AWS_LAMBDA_RUNTIME_API}/2018-06-01/runtime"
while true; do
# Block until next invocation arrives — response includes event payload + request ID header
RESPONSE=$(curl -sS -D /tmp/headers.txt "${RUNTIME_API}/invocation/next")
REQUEST_ID=$(grep -i "Lambda-Runtime-Aws-Request-Id" /tmp/headers.txt \
| tr -d '\r' | awk -F': ' '{print $2}')
echo "$RESPONSE" > /tmp/event.json
if RESULT=$(python3 /var/task/handler.py 2>/tmp/error.txt); then
curl -sS -X POST "${RUNTIME_API}/invocation/${REQUEST_ID}/response" \
-H "Content-Type: application/json" -d "$RESULT"
else
ERROR_MSG=$(head -1 /tmp/error.txt)
curl -sS -X POST "${RUNTIME_API}/invocation/${REQUEST_ID}/error" \
-H "Content-Type: application/json" \
-d "{\"errorMessage\":\"${ERROR_MSG}\",\"errorType\":\"HandlerError\"}"
fi
done
Bun runtime layer: fastest cold start in the custom runtime tier. Bun's JavaScriptCore engine initializes faster than Node.js's V8, and Bun natively parses TypeScript without a build step. The Bun binary for linux-x64 is ~40 MB unzipped — well within the 250 MB unzipped total limit. Cold start for a minimal MCP tool is 80–200 ms, compared to 200–400 ms for managed Node.js 22. The bootstrap script for a Bun layer is simple:
#!/bin/sh
# Bun custom runtime bootstrap
export PATH="/opt/bin:$PATH"
exec /opt/bin/bun /var/task/handler.ts
# Bun reads handler.ts natively — no tsc or esbuild step in the function ZIP
Deno runtime layer: security-first with explicit permissions. Deno's value proposition is its security model: every capability (network access, file reads, environment variables) requires an explicit flag at startup. The bootstrap script must list all required permissions or the handler will fail with PermissionDenied at runtime. The Deno binary is ~90 MB uncompressed, which typically produces a layer ZIP of 45–55 MB — above the 50 MB direct upload limit. Upload via S3 instead:
# Deno bootstrap with explicit permissions for an MCP tool
#!/bin/sh
export PATH="/opt/bin:$PATH"
exec /opt/bin/deno run \
--allow-env \
--allow-net \
--allow-read=/var/task \
/var/task/handler.ts
# --allow-net is required for AWS SDK calls (Bedrock, S3, DynamoDB, etc.)
# Scope further: --allow-net=bedrock-runtime.us-east-1.amazonaws.com for defense-in-depth
# Publish via S3 to bypass the 50 MB direct upload limit
aws s3 cp deno-runtime-layer.zip s3://my-artifact-bucket/layers/deno-runtime-layer.zip
aws lambda publish-layer-version \
--layer-name deno-runtime \
--content "S3Bucket=my-artifact-bucket,S3Key=layers/deno-runtime-layer.zip" \
--compatible-runtimes provided.al2023 \
--compatible-architectures x86_64 \
--region us-east-1
Cold start comparison across runtimes. The following table summarizes approximate cold-start ranges for a minimal MCP tool handler with no database connections, measured at 512 MB function memory:
| Runtime | Type | Cold start (approx.) | Layer size (unzipped) | Key trade-off |
|---|---|---|---|---|
| Node.js 22 | Managed | 200–400 ms | No layer needed | Largest ecosystem; V8; good baseline for most MCP tools |
| Bun 1.1.x | Custom runtime layer | 80–200 ms | ~40 MB | Fastest cold start; native TS; Bun.S3Client replaces AWS SDK S3 |
| Deno 2.x | Custom runtime layer | 150–350 ms | ~90 MB | Security model; no node_modules; S3 upload required |
| Python 3.12 | Managed | 100–250 ms | No layer needed | Fast; large stdlib; best for data-processing MCP tools |
| Python 3.13 | Custom runtime layer | 200–450 ms | ~70 MB | Slightly slower than managed 3.12; use only for 3.13-specific features |
Architecture-specific builds are mandatory. Bun and Deno publish separate binaries for linux-x64 (for x86_64 Lambda functions) and linux-aarch64 (for arm64 / Graviton3 functions). Deploying an x86_64 binary to an arm64 function returns Exec format error immediately at cold start. Build and publish separate layer versions for each architecture, and match the function's --architectures flag to the correct layer. Graviton3 functions cost ~20% less per invocation and are well-suited for MCP tools that are network-bound rather than CPU-bound.
Provisioned Concurrency for cold-start-sensitive MCP tools. Custom runtimes add 100–300 ms of cold-start overhead compared to their equivalent managed runtimes, because Lambda must load the runtime binary from the layer before executing bootstrap. For MCP tools that sit in a tight agent loop where every round-trip adds to user-perceived latency, a cold start that pushes total response time above the function timeout causes intermittent errors. Configure Provisioned Concurrency on 2–5 instances to keep execution environments pre-initialized. AliveMCP's 60-second probes show the bimodal response-time distribution in the Author tier's history view — warm invocations cluster at 50–100 ms while cold starts cluster at 400–900 ms, making it immediately obvious whether the timeout setting has enough headroom and whether cold starts are frequent enough to justify the provisioned concurrency cost (~$2–4/month per 512 MB function with 2 provisioned instances in us-east-1).
Consolidated failure modes
| Symptom | Root cause | Fix |
|---|---|---|
Cannot find module 'zod' at runtime despite layer being attached | Layer ZIP has node_modules/ at the root instead of nodejs/node_modules/; or layer's compatibleRuntimes does not match the function's runtime; or function is on a different architecture than the layer's compatibleArchitectures | Verify ZIP structure with unzip -l layer.zip | head -5 — first entry must be nodejs/; check compatibleRuntimes with aws lambda get-layer-version --query 'CompatibleRuntimes'; rebuild and republish if incorrect |
| Function ZIP is still 20–30 MB after adding the layer; cold start unchanged | esbuild --external flags omitted; esbuild bundles the layer packages into the function ZIP as well as having them in the layer — both copies exist, the function ZIP size is not reduced | Add --external:zod --external:"@aws-sdk/*" --external:jose to esbuild; verify no node_modules in function ZIP with unzip -l function.zip | grep node_modules — should return empty |
ResourceNotFoundException: Layer version does not exist on cold start or update-function-configuration | Layer version was deleted — either manually or via CDK RemovalPolicy.DESTROY during a stack update — while functions still referenced the ARN; Lambda enforces no referential integrity | Re-publish the layer from the same ZIP if available (creates a new version number); update functions to the new ARN; going forward set RemovalPolicy.RETAIN in CDK and RetentionPolicy: Retain in SAM for all layer constructs |
| Layer update not picked up — function keeps using the old version after new version was published | Lambda never auto-updates a function's layer ARN reference; publishing a new layer version has no effect on running functions until you explicitly call update-function-configuration with the new ARN | Run aws lambda update-function-configuration --function-name $fn --layers $NEW_LAYER_ARN for each function; CDK handles this automatically on cdk deploy when the layer asset hash changes |
| All MCP tool functions fail simultaneously after a "routine" dependency bump | Shared layer was promoted to all functions at once without canary testing; a breaking change in the new dependency version caused all tools to fail together; single-point-of-failure of a shared layer without blast-radius control | The canary workflow prevents this: always update one low-traffic function first, smoke-test it (check for errorMessage key in response body, not just HTTP status), then batch-update the remaining functions only if the canary is green |
AccessDeniedException: lambda:GetLayerVersion when attaching a cross-account layer | The consumer account (or the principal making the update-function-configuration call) is not granted in the layer version's resource-based policy in the platform account; every new layer version starts with an empty policy — permissions from previous versions do not carry over | In the platform account run add-layer-version-permission --version-number N --principal CONSUMER_ACCOUNT_ID or --principal "*" --principal-org-id o-xxx; verify with get-layer-version-policy |
Layer works in us-east-1 but functions in eu-west-1 cannot attach it | Lambda layer ARNs are region-scoped; a layer in us-east-1 cannot be referenced by a function in any other region; the cross-account permission grant must also be replicated per-region | Run publish-layer-version and add-layer-version-permission in every region your MCP platform operates in using a shell loop over region names; use the same ZIP for all regions for pure-JS packages |
Revocation breaks some function invocations but not others — intermittent AccessDeniedException | Permission check happens only at cold start when Lambda fetches the layer content; warm execution environments already have the layer loaded and continue working; after revocation, next cold start fails with AccessDeniedException | Re-add permission immediately to restore full availability; coordinate the next revocation with a 2–4 week deprecation window and confirm all consumer functions have migrated before revoking |
Runtime.InvalidEntrypoint — Lambda cannot execute bootstrap | bootstrap file is not at the root of the layer ZIP (e.g., accidentally placed in bin/bootstrap instead of the ZIP root, which maps to /opt/bin/bootstrap not /opt/bootstrap); or the file is not marked executable before zipping | Verify path with unzip -l layer.zip | grep bootstrap — should show bootstrap not bin/bootstrap; re-zip from the layer directory root with cd layer-dir && chmod +x bootstrap && zip -r9 ../layer.zip . |
| Cold starts time out; warm invocations succeed — only on custom runtime layer functions | Runtime binary initialization (Bun JIT warmup, Deno V8 startup, CPython import phase) exceeds the function timeout; cold start overhead scales with binary size — Deno's 90 MB binary adds more latency than Bun's 40 MB; a function with a 3-second timeout and a 2-second Deno cold start has no headroom for handler execution | Increase function timeout to give cold starts headroom (3 s minimum for Bun, 5 s for Deno, 3 s for CPython 3.13); add Provisioned Concurrency at 2–5 instances for latency-critical MCP tools in tight agent loops; monitor cold-start frequency with AliveMCP's response-time history in the Author tier |
AliveMCP and Lambda-based MCP server monitoring
When your MCP server runs as a fleet of Lambda functions behind a Function URL or API Gateway, a standard HTTP monitor that checks for a non-5xx status code is insufficient. Lambda error payloads are delivered as HTTP 200 with a body containing errorMessage and errorType — a function that fails with Cannot find module 'zod' returns HTTP 200, not HTTP 500. This pattern occurs precisely when a layer is misconfigured: wrong ZIP structure, missing --external flags during the layer build, or a deleted layer version that causes ResourceNotFoundException at cold start. A monitor that only checks for HTTP 200 will report the function as healthy when every MCP tool invocation is silently failing.
AliveMCP probes your Lambda function endpoints with a test MCP tools/call JSON-RPC payload every 60 seconds and parses the response body to distinguish a valid MCP JSON-RPC response from a Lambda runtime error payload, an HTTP 5xx from the Lambda invoke infrastructure, and a cold-start timeout. When a layer update causes a function to fail, AliveMCP detects the first failing probe within 60 seconds — before the first CloudWatch alarm fires (typically requiring 3–5 data points over 3–5 minutes) and before any real user traffic exercises the broken function. For canary deployments, use AliveMCP's status API to automate the promotion gate: if the AliveMCP check for the canary function is green after 5 minutes, proceed with the batch update; if it is red, abort. This turns external probe data into a real production signal rather than a manual observation tool.