AWS CI/CD · 2026-10-08 · CodePipeline/CodeDeploy arc
AWS CI/CD for MCP Servers: CodePipeline V2, CodeDeploy Blue-Green, and CDK Pipelines
The single most disruptive mistake in a CodePipeline setup is creating a GitHub V2 source action against a CodeStar Connection that is still in PENDING state — the pipeline creates without error, but every execution attempt fails immediately at Source with a cryptic authorization message that looks like an IAM problem, when the real fix is a one-time manual OAuth flow in the AWS Console that turns the connection to AVAILABLE. For MCP server deployments, CodePipeline solves the "push to GitHub and have a tested revision running in ECS within minutes" problem — a V2 pipeline with a CodeStar connection source, a CodeBuild build stage that produces an imagedefinitions.json artifact, and an ECS blue-green deploy action with CodeDeploy gives you atomic traffic shifting and a 60-minute rollback window with zero infrastructure management. The five topics that produce a production-ready CI/CD pipeline for MCP servers — CodePipeline V2 structure and GitHub integration (pipeline type V2, CodeStar Connection, artifact store, service role, execution modes, imagedefinitions.json, monorepo path filters), stage design patterns (action categories, runOrder, V2 pipeline variables, Manual Approval, staging-to-prod flow), CodeDeploy EC2 rolling deployments (CodeDeploy agent, appspec.yml structure, lifecycle hooks, deployment configurations, automatic rollback), ECS blue-green deployments (dual target groups, traffic shifting policies, terminationWaitTimeInMinutes, BeforeAllowTraffic/AfterAllowTraffic hooks, ALB deregistration delay for SSE), and CDK Pipelines self-mutation (selfMutation, synth step, cross-account bootstrap, addStage, waves for parallel multi-region rollout) — each contains sharp edges that produce silent failures, permanently stuck deployments, or infrastructure drift that only surfaces after a failed rollback. The pipeline service role needs iam:PassRole scoped to the ECS task execution role with a condition limiting the pass to ecs-tasks.amazonaws.com — without that condition the role is effectively unrestricted and will fail an AWS Security Hub finding for overly permissive role-passing. ECS blue-green requires the service to have been created with deploymentController: { type: CODE_DEPLOY } — you cannot switch the deployment controller type in-place after service creation, so retrofitting blue-green onto an existing ECS service means destroying and recreating the service. This guide synthesizes all five topics into three structural patterns: pipeline setup and GitHub integration, deployment strategies (rolling EC2, ECS blue-green, approval gates), and CDK Pipelines infrastructure as code with self-mutation.
TL;DR
- Pipeline setup and GitHub integration: create a CodePipeline V2 pipeline (not V1 — V2 is required for pipeline variables and monorepo path filtering) with a GitHub V2 source action backed by a CodeStar Connection. The connection must reach
AVAILABLEstate via a one-time manual OAuth authorization in the AWS Console before any pipeline execution will succeed — a connection inPENDINGstate causes immediate Source failure that looks like an IAM error but is an authorization gap. Configure the artifact store as an S3 bucket with versioning enabled, a KMS CMK for encryption (kms:GenerateDataKeyandkms:Decrypton the pipeline service role), and a lifecycle rule expiring objects after 30 days and noncurrent versions after 7 days. The pipeline service role needscodestar-connections:UseConnection,codebuild:StartBuildandcodebuild:BatchGetBuilds,ecs:RegisterTaskDefinitionandecs:UpdateService, andiam:PassRolescoped to the ECS task execution role ARN with conditioniam:PassedToService: ecs-tasks.amazonaws.com. In CodeBuild, setOutputArtifactFormat: CODEBUILD_CLONE_REFon the Source action to give CodeBuild a full git clone with history for SHA-based image tags; the CodeBuild project role must also havecodestar-connections:UseConnectionfor this to work. Writeimagedefinitions.jsonin the CodeBuild post_build phase:printf '[{"name":"app","imageUri":"%s"}]' "$ECR_IMAGE" > imagedefinitions.json— the ECS deploy action reads this file to update the task definition; a missing or malformed file causes the deploy action to fail with no helpful error message. Use execution modeSUPERSEDED(new push cancels in-flight execution) for nearly all MCP deployments;PARALLELmode is dangerous for ECS deployments where two executions would update the same service simultaneously. - Deployment strategies — rolling EC2, ECS blue-green, approval gates: for EC2 deployments, the CodeDeploy agent must be installed and running on every instance before the first deployment — instances without the agent are silently skipped (zero failures reported, but old code keeps running). Put
appspec.ymlat the bundle root withfile_exists_behavior: OVERWRITEon each files entry (the defaultDISALLOWcauses BeforeInstall to fail on update deployments when destination files already exist). Use ValidateService with a retry loop —for i in $(seq 1 10); do curl -sf http://localhost:8080/health && exit 0; sleep 5; done; exit 1— non-zero exit triggers automatic rollback, so the loop must account for slow MCP server startup. For ECS blue-green, configure two target groups (blue receives production traffic on port 443, green receives the test listener on port 8443), setterminationWaitTimeInMinutes: 60to keep the blue environment alive for 60 minutes after full traffic shift (rollback is available viastop-deploymentduring this window), and increase the ALB deregistration delay from the default 300 seconds to 600 seconds on the blue target group to avoid dropping in-flight MCP SSE connections during traffic shift. For staging-to-prod flow, structure the pipeline as: Source → Build → DeployStaging → TestAndApprove (integration tests at runOrder 1 in parallel with Manual Approval at runOrder 1 — both must succeed before the stage advances) → DeployProd. V2 pipeline variables declared on the Build action namespace and exported from CodeBuild buildspec asexported-variablesare consumed downstream as#{namespace.variable}— enabling dynamic ECR image tag injection into the ECS deploy action without hardcoding. - CDK Pipelines self-mutation and cross-account deployment: import from
aws-cdk-lib/pipelines(NOTaws-codepipeline— these are completely different APIs). SetselfMutation: trueso the pipeline's first stage is alwaysUpdatePipeline— it updates the pipeline's own CloudFormation stack before deploying any application stage, meaning you can add new deployment stages in TypeScript, push, and the pipeline installs itself without a manual CloudFormation update. The synth step should use aCodeBuildStepwithinstallCommands: ['npm install -g aws-cdk@2'],commands: ['cd infrastructure && npm ci && npm run build && npx cdk synth'], andprimaryOutputDirectory: 'cdk.out'— the primary output directory must match the CDK output directory exactly or the pipeline will fail to find CloudFormation templates after synth. For cross-account deployment, runcdk bootstrap --trust <pipeline-account-id> --trust-for-lookup <pipeline-account-id> --cloudformation-execution-policies arn:aws:iam::aws:policy/AdministratorAccessin every target account — without--trust-for-lookup, VPC ID and AMI ID context lookups during synth fail with unresolved references. Usepipeline.addWave()with multipleaddStage()calls to deploy all stages in parallel for multi-region MCP rollout (us-east-1, eu-west-1, and ap-southeast-1 simultaneously). Avoid non-deterministic synth: never useDate.now()in logical IDs, pin the CDK version inpackage.json, and commitcdk.context.json— a pipeline that perpetually self-mutates blocks all application deployments indefinitely.
Pattern 1 — Pipeline Setup and GitHub Integration
CodePipeline V2 and the CodeStar Connection authorization gap
CodePipeline has two pipeline types: V1 (legacy) and V2 (current). Use V2 for all new MCP server pipelines — V2 adds pipeline-level variables that can be declared on one action and consumed in downstream stages, and adds trigger filters including push.filePaths.includes for monorepo path filtering so pushes touching only docs/ do not trigger a full rebuild and deploy. The V1 type will not receive new features and AWS has been quietly guiding customers to V2 since 2023.
The GitHub V2 source action uses CodeStar Connections, which is the correct mechanism for all new GitHub integrations (the older GitHub V1 action used OAuth tokens stored in Secrets Manager and is deprecated). The critical property of a CodeStar Connection that trips nearly every first-time setup: a newly created connection starts in PENDING state. The pipeline creates successfully, the source action saves without error — but every pipeline execution attempt fails immediately at the Source stage because the connection has not been authorized. The fix is a one-time manual step: navigate to the CodeStar Connections console, select the connection, click "Update pending connection," and complete the GitHub OAuth flow. Once completed, the connection reaches AVAILABLE state and all future executions succeed. There is no AWS CLI or CloudFormation mechanism to complete this OAuth step — it must be done by a human in the console before the pipeline is used in production.
# Create a CodeStar Connection for GitHub
aws codestar-connections create-connection \
--provider-type GitHub \
--connection-name mcp-server-github \
--region us-east-1
# Returns: ConnectionArn — status will be PENDING
# After manually authorizing in the console, verify AVAILABLE state:
aws codestar-connections get-connection \
--connection-arn arn:aws:codestar-connections:us-east-1:123456789012:connection/abc123 \
--query 'Connection.ConnectionStatus' \
--output text
# Must return: AVAILABLE (not PENDING)
# CloudFormation: V2 pipeline with GitHub V2 source
# PipelineType: V2 is required for variables and path filters
Resources:
McpPipeline:
Type: AWS::CodePipeline::Pipeline
Properties:
PipelineType: V2
RoleArn: !GetAtt PipelineRole.Arn
ArtifactStore:
Type: S3
Location: !Ref ArtifactBucket
EncryptionKey:
Type: KMS
Id: !Ref ArtifactKMSKey
Triggers:
- ProviderType: CodeStarSourceConnection
GitConfiguration:
SourceActionName: GitHub
Push:
- Branches:
Includes: [main]
FilePaths:
Includes: [src/**, infrastructure/**]
# Unmatched pushes (e.g. docs-only) are silently skipped
Stages:
- Name: Source
Actions:
- Name: GitHub
ActionTypeId:
Category: Source
Owner: AWS
Provider: CodeStarSourceConnection
Version: "1"
Configuration:
ConnectionArn: !Ref GitHubConnection
FullRepositoryId: myorg/mcp-server
BranchName: main
OutputArtifactFormat: CODEBUILD_CLONE_REF
# Full git clone with history for SHA-based image tags
# Requires CodeBuild role to also have codestar-connections:UseConnection
OutputArtifacts:
- Name: SourceArtifact
- Name: Build
Actions:
- Name: BuildAndPush
ActionTypeId:
Category: Build
Owner: AWS
Provider: CodeBuild
Version: "1"
Namespace: BuildVars
Configuration:
ProjectName: !Ref CodeBuildProject
InputArtifacts:
- Name: SourceArtifact
OutputArtifacts:
- Name: BuildArtifact
Artifact store and pipeline service role
The artifact store is an S3 bucket that CodePipeline uses to pass artifacts between stages. Every artifact — source code ZIP, build output, appspec, imagedefinitions.json — passes through this bucket. Three configuration requirements matter for a production MCP server pipeline: versioning must be enabled (CodePipeline requires it and will fail to create without it), encryption should use a KMS CMK so you can audit key usage and rotate the key independently of the bucket, and a lifecycle rule should expire objects after 30 days (noncurrent versions after 7 days) to prevent unbounded storage growth from every push ever made to the branch.
The pipeline service role requires precise permissions. The most common gap is iam:PassRole — when the ECS deploy action registers a new task definition and updates the ECS service, it passes the ECS task execution role to ECS. The pipeline service role must have iam:PassRole on that task execution role ARN. Scope the permission with the condition iam:PassedToService: ecs-tasks.amazonaws.com to prevent the role from being used to pass any role to any service — without this condition, a broadly scoped iam:PassRole is a privilege escalation vector.
# Pipeline service role inline policy
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "CodeStarConnection",
"Effect": "Allow",
"Action": "codestar-connections:UseConnection",
"Resource": "arn:aws:codestar-connections:us-east-1:123456789012:connection/*"
},
{
"Sid": "CodeBuild",
"Effect": "Allow",
"Action": ["codebuild:StartBuild", "codebuild:BatchGetBuilds", "codebuild:StopBuild"],
"Resource": "arn:aws:codebuild:us-east-1:123456789012:project/mcp-server-build"
},
{
"Sid": "ECSDeployRead",
"Effect": "Allow",
"Action": [
"ecs:RegisterTaskDefinition",
"ecs:DescribeTaskDefinition",
"ecs:DescribeServices"
],
"Resource": "*"
},
{
"Sid": "ECSDeployWrite",
"Effect": "Allow",
"Action": ["ecs:UpdateService"],
"Resource": "arn:aws:ecs:us-east-1:123456789012:service/mcp-cluster/mcp-server"
},
{
"Sid": "PassRoleToECS",
"Effect": "Allow",
"Action": "iam:PassRole",
"Resource": "arn:aws:iam::123456789012:role/McpTaskExecutionRole",
"Condition": {
"StringEquals": { "iam:PassedToService": "ecs-tasks.amazonaws.com" }
}
},
{
"Sid": "ArtifactStoreKMS",
"Effect": "Allow",
"Action": ["kms:GenerateDataKey", "kms:Decrypt"],
"Resource": "arn:aws:kms:us-east-1:123456789012:key/artifact-key-id"
},
{
"Sid": "ArtifactStoreS3",
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject", "s3:GetObjectVersion"],
"Resource": "arn:aws:s3:::mcp-pipeline-artifacts/*"
},
{
"Sid": "SNSApproval",
"Effect": "Allow",
"Action": "sns:Publish",
"Resource": "arn:aws:sns:us-east-1:123456789012:mcp-deploy-approvals"
}
]
}
CodeBuild stage and imagedefinitions.json
The build stage is where Docker images are built, tagged with the git commit SHA, pushed to ECR, and the deployment artifact is assembled. The commit SHA as the image tag is the right approach for MCP server images: it makes every image version traceable to a specific commit, prevents the "latest" ambiguity where two environments might run images with the same tag but different code, and simplifies rollback (revert the tag in the task definition).
The imagedefinitions.json file is the bridge between CodeBuild and the ECS deploy action. It maps container names to ECR image URIs and is written in the post_build phase. A missing file, a file in the wrong output artifact, or a file with mismatched container name causes the ECS deploy action to fail with a generic "no container mapping found" error that does not indicate which artifact or container name is wrong — so validate both the container name (it must exactly match the container name in the ECS task definition, case-sensitive) and that the file is included in the CodeBuild output artifact.
# buildspec.yml for MCP server build stage
version: 0.2
env:
variables:
REPO_NAME: mcp-server
AWS_DEFAULT_REGION: us-east-1
exported-variables:
- IMAGE_TAG
- ECR_IMAGE
phases:
pre_build:
commands:
- echo Logging in to Amazon ECR...
- aws ecr get-login-password --region $AWS_DEFAULT_REGION | docker login --username AWS --password-stdin $AWS_ACCOUNT_ID.dkr.ecr.$AWS_DEFAULT_REGION.amazonaws.com
- COMMIT_SHA=$(git rev-parse --short HEAD)
- IMAGE_TAG="${COMMIT_SHA}-${CODEBUILD_BUILD_NUMBER}"
- ECR_REPO="$AWS_ACCOUNT_ID.dkr.ecr.$AWS_DEFAULT_REGION.amazonaws.com/$REPO_NAME"
- ECR_IMAGE="$ECR_REPO:$IMAGE_TAG"
build:
commands:
- echo Building Docker image at $(date)
- docker build -t $ECR_IMAGE -f Dockerfile .
- docker tag $ECR_IMAGE $ECR_REPO:latest
post_build:
commands:
- echo Pushing Docker image...
- docker push $ECR_IMAGE
- docker push $ECR_REPO:latest
# imagedefinitions.json: container name MUST match ECS task definition container name exactly
- printf '[{"name":"mcp-server","imageUri":"%s"}]' "$ECR_IMAGE" > imagedefinitions.json
# For multiple containers in the same task:
# printf '[{"name":"mcp-server","imageUri":"%s"},{"name":"sidecar","imageUri":"%s"}]' "$ECR_IMAGE" "$SIDECAR_IMAGE" > imagedefinitions.json
artifacts:
files:
- imagedefinitions.json
- appspec.yaml # required for ECS CodeDeploy blue-green
- taskdef.json # required for ECS CodeDeploy blue-green
discard-paths: yes
For ECS blue-green deployments via CodeDeploy, the build artifact must include three files: imagedefinitions.json (for standard ECS rolling), appspec.yaml (CodeDeploy deployment specification), and taskdef.json (the ECS task definition JSON with a <IMAGE1_NAME> placeholder that CodeDeploy replaces). Use jq in the post_build phase to inject the new image URI into the task definition: jq --arg IMAGE "$ECR_IMAGE" '.containerDefinitions[0].image = $IMAGE' taskdef.template.json > taskdef.json. The appspec.yaml for ECS should contain the <TASK_DEFINITION> placeholder that instructs CodeDeploy to register the taskdef.json from the artifact as a new task revision.
Execution modes and monorepo path filtering
CodePipeline V2 supports three execution modes. SUPERSEDED means a new pipeline execution cancels any in-flight execution of the same pipeline — this is the correct mode for most MCP server deployments because you only care about shipping the latest commit, not every intermediate commit. QUEUED means every commit is deployed in order, with each execution waiting for the previous to complete — useful when each commit represents a database migration or other ordered operation. PARALLEL allows multiple simultaneous executions of the same pipeline — never use this for ECS services where two parallel executions would call UpdateService on the same cluster and service simultaneously; the second call wins but the first's deployment may be partially applied, leaving the service in an inconsistent state.
The V2 trigger filter push.filePaths.includes enables monorepo path filtering: pushes that touch no matching path are silently skipped — no execution is created, no failure is recorded. This is correct behavior (unmatched pushes should not trigger the pipeline) but can be confusing when debugging: if a push does not trigger the pipeline, verify the changed file paths against the filter configuration rather than assuming an IAM or webhook problem.
Pattern 2 — Deployment Strategies: Rolling EC2, ECS Blue-Green, Approval Gates
CodeDeploy EC2 rolling deployments
The foundational requirement for CodeDeploy EC2 deployments is the CodeDeploy agent running on every instance in the deployment group. The agent is a daemon process that polls the CodeDeploy service for deployment instructions and executes lifecycle hooks locally. New Auto Scaling Group instances do not have the agent unless it is baked into the AMI or installed via an ASG lifecycle hook using SSM Run Command with the AWS-ConfigureAWSPackage document. The failure mode is silent: instances without the agent are skipped, the deployment reports zero failures and zero successes for those instances, and the old code revision keeps running. This looks like a successful deployment from the CodePipeline perspective.
# Install CodeDeploy agent via SSM Run Command (for existing instances)
aws ssm send-command \
--document-name "AWS-ConfigureAWSPackage" \
--parameters '{"action":["Install"],"name":["AWSCodeDeployAgent"]}' \
--targets '[{"Key":"tag:Env","Values":["production"]}]' \
--region us-east-1
# appspec.yml — must be at the bundle root (same directory level as your app files)
version: 0.0
os: linux
files:
- source: /
destination: /opt/mcp-server
file_exists_behavior: OVERWRITE
# OVERWRITE required — default DISALLOW refuses to overwrite existing files
# on update deployments and causes BeforeInstall to fail
permissions:
- object: /opt/mcp-server
owner: mcp
group: mcp
mode: "755"
type:
- directory
- object: /opt/mcp-server
pattern: "**"
owner: mcp
group: mcp
mode: "644"
type:
- file
hooks:
BeforeInstall:
- location: scripts/stop_server.sh
timeout: 60
runas: root
AfterInstall:
- location: scripts/install_deps.sh
timeout: 120
runas: mcp
ApplicationStart:
- location: scripts/start_server.sh
timeout: 30
runas: root
ValidateService:
- location: scripts/validate_service.sh
timeout: 120
runas: root
The lifecycle hook scripts define the deployment choreography. The BeforeInstall hook should stop the running MCP server gracefully — send SIGTERM first (which triggers graceful drain of open SSE connections) and fall back to SIGKILL after a timeout. The AfterInstall hook installs or updates dependencies (e.g., npm ci --production or pip install -r requirements.txt). ApplicationStart starts the service (typically by enabling and starting a systemd unit). ValidateService is the most important hook: it must verify the new revision is actually serving requests before CodeDeploy considers the deployment successful.
#!/bin/bash
# scripts/stop_server.sh
set -e
systemctl stop mcp-server || true
# Wait for graceful shutdown of SSE connections (up to 30 seconds)
for i in $(seq 1 30); do
if ! systemctl is-active --quiet mcp-server; then
echo "Server stopped gracefully"
exit 0
fi
sleep 1
done
# Force stop if still running
systemctl kill -s SIGKILL mcp-server || true
---
#!/bin/bash
# scripts/validate_service.sh
# Non-zero exit triggers automatic rollback
set -e
MAX_RETRIES=10
HEALTH_URL="http://localhost:8080/health"
for i in $(seq 1 $MAX_RETRIES); do
if curl -sf "$HEALTH_URL" --max-time 5; then
echo "Health check passed on attempt $i"
exit 0
fi
echo "Attempt $i failed, retrying in 5s..."
sleep 5
done
echo "Health check failed after $MAX_RETRIES attempts"
exit 1
Deployment configurations control how many instances are updated simultaneously. CodeDeployDefault.OneAtATime keeps all minus one instance healthy throughout the deployment (safest for production — never takes more than one instance offline), CodeDeployDefault.HalfAtATime takes down 50% at a time for a two-round rollout, and CodeDeployDefault.AllAtOnce takes all instances offline simultaneously — never use AllAtOnce in production because a failed deployment leaves zero healthy instances serving traffic. Enable automatic rollback on DEPLOYMENT_FAILURE and DEPLOYMENT_STOP_ON_ALARM with a CloudWatch alarm that fires when the application error rate exceeds a threshold — this gives you automated rollback when the ValidateService hook passes but the new revision develops errors under real traffic.
ECS blue-green deployments with CodeDeploy
ECS blue-green is the deployment strategy for production MCP servers on ECS: the new task revision starts in a parallel "green" environment behind an ALB test listener, traffic shifts from the old "blue" environment to green only after health checks pass, and the blue environment remains live for a configurable window in case rollback is needed. The constraint that catches teams by surprise is that the deployment controller type on an ECS service is immutable — a service created with deploymentController: { type: ECS }` (the default) cannot be changed to `CODE_DEPLOY without recreating the service. Plan the deployment controller at service creation time.
# ECS service requiring CODE_DEPLOY deployment controller for blue-green
aws ecs create-service \
--cluster mcp-cluster \
--service-name mcp-server \
--task-definition mcp-server:1 \
--desired-count 2 \
--launch-type FARGATE \
--deployment-controller type=CODE_DEPLOY \
--network-configuration "awsvpcConfiguration={subnets=[subnet-abc,subnet-def],securityGroups=[sg-123],assignPublicIp=DISABLED}" \
--load-balancers '[
{
"targetGroupArn": "arn:aws:elasticloadbalancing:us-east-1:123456789012:targetgroup/mcp-blue/abc123",
"containerName": "mcp-server",
"containerPort": 8080
}
]'
# appspec.yaml for ECS blue-green deployment
# Note: .yaml extension required (not .yml) for ECS CodeDeploy
version: 0.0
Resources:
- TargetService:
Type: AWS::ECS::Service
Properties:
TaskDefinition:
# is a literal placeholder — CodeDeploy replaces it
# with the ARN of the task revision registered from taskdef.json
LoadBalancerInfo:
ContainerName: mcp-server
ContainerPort: 8080
PlatformVersion: LATEST
Hooks:
- BeforeAllowTraffic: McpSmokeTestFunction
# Lambda ARN or function name — called after green passes health checks
# but before any production traffic shifts to green
# Must call codedeploy:PutLifecycleEventHookExecutionStatus
- AfterAllowTraffic: McpPostDeployFunction
# CodeDeploy deployment group for ECS blue-green
aws deploy create-deployment-group \
--application-name mcp-server-app \
--deployment-group-name mcp-server-prod \
--service-role-arn arn:aws:iam::123456789012:role/CodeDeployECSRole \
--deployment-config-name CodeDeployDefault.ECSLinear10PercentEvery1Minutes \
--ecs-services clusterName=mcp-cluster,serviceName=mcp-server \
--load-balancer-info '{
"targetGroupPairInfoList": [{
"targetGroups": [
{"name": "mcp-blue"},
{"name": "mcp-green"}
],
"prodTrafficRoute": {"listenerArns": ["arn:aws:elasticloadbalancing:...:listener/app/mcp-alb/.../prod"]},
"testTrafficRoute": {"listenerArns": ["arn:aws:elasticloadbalancing:...:listener/app/mcp-alb/.../test"]}
}]
}' \
--blue-green-deployment-configuration '{
"terminateBlueInstancesOnDeploymentSuccess": {
"action": "TERMINATE",
"terminationWaitTimeInMinutes": 60
},
"deploymentReadyOption": {
"actionOnTimeout": "CONTINUE_DEPLOYMENT",
"waitTimeInMinutes": 0
}
}'
Traffic shifting policies determine how quickly production traffic moves from blue to green. ECSAllAtOnce shifts all traffic immediately — use only for staging environments. ECSLinear10PercentEvery1Minutes shifts 10% per minute for a 10-minute gradual shift while monitoring error rates — the right choice for production MCP servers where client reconnection on SSE drop is noticeable. ECSCanary10Percent5Minutes sends 10% to green for 5 minutes then shifts the remaining 90% — good for quick validation before full shift. ECSCanary10Percent15Minutes holds 10% longer for deeper canary observation when the MCP server calls downstream services with latency that only appears at sustained load.
The terminationWaitTimeInMinutes: 60 setting is critical for MCP servers. After the full traffic shift completes, blue tasks remain alive and registered in the blue target group for 60 minutes. During this window you can trigger a rollback via aws deploy stop-deployment without re-running the full deployment pipeline. After the window closes, blue tasks are terminated and rollback requires a new full deployment from source. Set the ALB deregistration delay on the blue target group to at least 600 seconds (10 minutes) to avoid dropping long-running MCP SSE connections during the initial deregistration phase at traffic shift start.
# BeforeAllowTraffic Lambda hook
# This Lambda runs after green passes health checks, before any traffic shifts
import boto3
import json
codedeploy = boto3.client('codedeploy')
def lambda_handler(event, context):
deployment_id = event['DeploymentId']
lifecycle_event_hook_execution_id = event['LifecycleEventHookExecutionId']
try:
# Run smoke tests against the test listener (green environment)
# The test listener routes to green; production listener still routes to blue
test_alb_endpoint = "https://mcp-test.internal:8443"
run_smoke_tests(test_alb_endpoint)
status = 'Succeeded'
except Exception as e:
print(f"Smoke test failed: {e}")
status = 'Failed'
finally:
# CRITICAL: always call PutLifecycleEventHookExecutionStatus
# Failing to call this causes the deployment to stall indefinitely
codedeploy.put_lifecycle_event_hook_execution_status(
deploymentId=deployment_id,
lifecycleEventHookExecutionId=lifecycle_event_hook_execution_id,
status=status
)
def run_smoke_tests(endpoint):
import urllib.request
req = urllib.request.urlopen(f"{endpoint}/health", timeout=10)
if req.status != 200:
raise Exception(f"Health check returned {req.status}")
# Test that the MCP server responds to tool list requests
req2 = urllib.request.urlopen(f"{endpoint}/mcp/tools/list", timeout=10)
if req2.status != 200:
raise Exception(f"Tool list returned {req2.status}")
Stage design and approval gates
A production stage design for MCP servers typically follows: Source → Build → DeployStaging → TestAndApprove → DeployProd. The TestAndApprove stage uses runOrder to run integration tests in parallel with a Manual Approval action — both actions have runOrder: 1, and the stage does not advance until both succeed. This means an approver cannot accidentally approve a deployment that has failing integration tests (the stage waits for both to succeed simultaneously, not either independently).
# TestAndApprove stage with parallel integration tests and manual approval
- Name: TestAndApprove
Actions:
- Name: IntegrationTests
RunOrder: 1
ActionTypeId:
Category: Test
Owner: AWS
Provider: CodeBuild
Version: "1"
Configuration:
ProjectName: mcp-integration-tests
InputArtifacts:
- Name: BuildArtifact
- Name: ManualApproval
RunOrder: 1
ActionTypeId:
Category: Approval
Owner: AWS
Provider: Manual
Version: "1"
Configuration:
NotificationArn: arn:aws:sns:us-east-1:123456789012:mcp-deploy-approvals
CustomData: "Deploy #{BuildVars.IMAGE_TAG} to production. Integration test results at #{BuildVars.TEST_REPORT_URL}"
ExternalEntityLink: "https://github.com/myorg/mcp-server/commit/#{BuildVars.COMMIT_SHA}"
# #{BuildVars.IMAGE_TAG} references V2 pipeline variable exported from Build stage
# SNS topic policy must allow codepipeline.amazonaws.com to publish
- Name: DeployProd
Actions:
- Name: BlueGreenDeploy
RunOrder: 1
ActionTypeId:
Category: Deploy
Owner: AWS
Provider: CodeDeployToECS
Version: "1"
Configuration:
ApplicationName: mcp-server-app
DeploymentGroupName: mcp-server-prod
TaskDefinitionTemplateArtifact: BuildArtifact
TaskDefinitionTemplatePath: taskdef.json
AppSpecTemplateArtifact: BuildArtifact
AppSpecTemplatePath: appspec.yaml
Image1ArtifactName: BuildArtifact
Image1ContainerName: IMAGE1_NAME
InputArtifacts:
- Name: BuildArtifact
The Manual Approval action's SNS topic must have a topic policy granting codepipeline.amazonaws.com permission to publish — without this, the approval notification is silently dropped but the pipeline still waits indefinitely for manual approval, which looks like a hung pipeline rather than a missing notification. Approval tokens expire after 7 days — if no action is taken within 7 days, the execution fails. For teams using Slack, AWS Chatbot integration can deliver approval requests to a Slack channel with one-click approve/reject buttons.
V2 pipeline variables declared via the Namespace property on an action and exported from CodeBuild via exported-variables in the buildspec are consumed downstream as #{namespace.variable}. A common failure is exporting a variable name in exported-variables that does not appear in the buildspec env.variables or env.exported-variables section — the variable resolves to an empty string downstream with no warning, and the deploy action receives a malformed image tag.
Pattern 3 — Infrastructure as Code: CDK Pipelines Self-Mutation
CDK Pipelines construct and selfMutation
CDK Pipelines is an L3 construct for defining CI/CD pipelines that deploy CDK applications. Import it from aws-cdk-lib/pipelines — not from aws-codepipeline. These are entirely separate APIs: aws-cdk-lib/pipelines is the high-level CDK Pipelines construct that manages CodePipeline, CodeBuild, and deployment stages as a single abstraction; aws-codepipeline is the low-level L2 construct for building raw CodePipeline resources. Using the wrong import produces confusing type errors that do not indicate the root cause.
// CDK Pipelines pipeline definition
// File: infrastructure/lib/pipeline-stack.ts
import { Stack, StackProps } from 'aws-cdk-lib';
import { Construct } from 'constructs';
import { CodePipeline, CodePipelineSource, CodeBuildStep, ManualApprovalStep } from 'aws-cdk-lib/pipelines';
// ^^ import from aws-cdk-lib/pipelines — NOT aws-codepipeline
import * as codebuild from 'aws-cdk-lib/aws-codebuild';
import { McpServerStage } from './mcp-server-stage';
export class PipelineStack extends Stack {
constructor(scope: Construct, id: string, props?: StackProps) {
super(scope, id, props);
const pipeline = new CodePipeline(this, 'McpPipeline', {
pipelineName: 'mcp-server-pipeline',
selfMutation: true,
// selfMutation: true adds UpdatePipeline as the first stage
// after every push, pipeline updates its own CloudFormation stack
// before deploying any application stage — add stages in TypeScript,
// push, and the pipeline installs itself without manual cfn update
crossAccountKeys: true,
// required for cross-account deployments — creates KMS key
// shared with target accounts for artifact encryption
dockerEnabledForSynth: true,
dockerEnabledForSelfMutation: true,
// required if synth step builds Docker images or if app uses
// ContainerImage.fromAsset() — enables Docker daemon in CodeBuild
synth: new CodeBuildStep('Synth', {
input: CodePipelineSource.connection('myorg/mcp-server', 'main', {
connectionArn: 'arn:aws:codestar-connections:us-east-1:123456789012:connection/abc123',
triggerOnPush: true,
}),
installCommands: [
'npm install -g aws-cdk@2.100.0',
// Pin CDK version to prevent self-mutation from unpinned package hashes
],
commands: [
'cd infrastructure',
'npm ci',
'npm run build',
'npx cdk synth',
],
primaryOutputDirectory: 'infrastructure/cdk.out',
// primaryOutputDirectory must match the CDK output directory
// mismatch causes pipeline to fail looking for CloudFormation templates
buildEnvironment: {
buildImage: codebuild.LinuxBuildImage.STANDARD_7_0,
privileged: true, // required for dockerEnabledForSynth
},
}),
});
// Staging deployment
const stagingStage = pipeline.addStage(
new McpServerStage(this, 'Staging', {
env: { account: '111111111111', region: 'us-east-1' },
})
);
// Post-stage integration tests using CloudFormation outputs
stagingStage.addPost(
new CodeBuildStep('IntegrationTests', {
commands: [
'npm ci',
'npx jest --testPathPattern=integration --testTimeout=30000',
],
envFromCfnOutputs: {
// ALB_DNS_NAME comes from a CfnOutput in McpServerStack
ALB_DNS_NAME: stagingStage.stackOutputs.albDnsName,
},
// envFromCfnOutputs maps CloudFormation stack outputs to env vars
// in this CodeBuild step — no secondary DynamoDB lookup required
})
);
// Production gate
stagingStage.addPost(new ManualApprovalStep('ApproveProduction'));
// Production wave — multi-region parallel deployment
const prodWave = pipeline.addWave('Production');
for (const region of ['us-east-1', 'eu-west-1', 'ap-southeast-1']) {
prodWave.addStage(
new McpServerStage(this, `Prod-${region}`, {
env: { account: '222222222222', region },
})
);
}
// All three region stages deploy in parallel within the Production wave
}
}
Cross-account bootstrap and context lookups
CDK Pipelines cross-account deployment requires the CDKToolkit bootstrap stack to be installed in every target account. The bootstrap stack creates a deploy role, a CloudFormation execution role, a lookup role, and an asset bucket — all with trust relationships back to the pipeline account. Run the bootstrap command in each target account with the correct trust flags:
# Bootstrap target account (run once per account-region combination)
# Must be run while authenticated to the TARGET account (not the pipeline account)
cdk bootstrap \
aws://222222222222/us-east-1 \
--trust 111111111111 \
--trust-for-lookup 111111111111 \
--cloudformation-execution-policies arn:aws:iam::aws:policy/AdministratorAccess
# --trust: allows pipeline account to assume the deploy role in this account
# --trust-for-lookup: allows pipeline account to assume the lookup role
# for VPC ID, AMI ID, and other context lookups during cdk synth
# WITHOUT --trust-for-lookup, synth fails with:
# "Cannot use 'Vpc.fromLookup()': Resolution failed: Context lookup failed"
# Also bootstrap the pipeline account itself (for asset publishing)
cdk bootstrap \
aws://111111111111/us-east-1 \
--cloudformation-execution-policies arn:aws:iam::aws:policy/AdministratorAccess
# Verify bootstrap version compatibility — CDK requires a minimum qualifier version
# If bootstrap version is too old, deployment fails with:
# "CDKToolkit stack version 6 is below the minimum required version 14"
aws cloudformation describe-stacks \
--stack-name CDKToolkit \
--region us-east-1 \
--profile target-account \
--query 'Stacks[0].Parameters[?ParameterKey==`BootstrapVersion`].ParameterValue' \
--output text
Avoiding the self-mutation loop
The self-mutation feature is powerful but has a failure mode that blocks all deployments: if the CDK synth output is non-deterministic, the pipeline perpetually detects that the pipeline stack needs updating, runs UpdatePipeline, and then restarts — never reaching the application deployment stages. The three common causes of non-deterministic synth are: using Date.now() or Math.random() in logical IDs or construct names (generates different values each synth), unpinned CDK versions in package.json (a patch release can change asset hashing behavior, producing different asset keys per synth), and uncommitted context lookups in cdk.context.json (VPC ID lookups performed during synth write to cdk.context.json and subsequent synths read from cache — but if the cache file is not committed, each synth performs a fresh lookup that may resolve to a slightly different context object).
// WRONG — non-deterministic logical ID causes self-mutation loop
export class McpServerStack extends Stack {
constructor(scope: Construct, id: string, props?: StackProps) {
super(scope, id, props);
// NEVER use Date.now() or Math.random() in construct IDs
new CfnOutput(this, `Output-${Date.now()}`, { value: 'foo' });
// ^^^^^^^^^^^^ different each synth = perpetual loop
}
}
// CORRECT — deterministic construct IDs
export class McpServerStack extends Stack {
readonly albDnsName: CfnOutput;
constructor(scope: Construct, id: string, props?: StackProps) {
super(scope, id, props);
const alb = new elbv2.ApplicationLoadBalancer(this, 'McpAlb', {
vpc,
internetFacing: true,
});
// Fixed logical ID — identical across every synth
this.albDnsName = new CfnOutput(this, 'AlbDnsName', {
value: alb.loadBalancerDnsName,
exportName: `${this.stackName}-AlbDnsName`,
});
}
}
// package.json — pin CDK version to prevent hash drift
{
"dependencies": {
"aws-cdk-lib": "2.100.0", // exact version, not ^2.100.0
"constructs": "10.3.0" // exact version
},
"devDependencies": {
"aws-cdk": "2.100.0" // exact version
}
}
# cdk.context.json — commit this file to the repository
# Context lookups (VPC IDs, AMI IDs, availability zones) are cached here
# Without committed context, each synth performs fresh lookups that may differ
{
"availability-zones:account=222222222222:region=us-east-1": [
"us-east-1a",
"us-east-1b",
"us-east-1c"
],
"vpc-provider:account=222222222222:region=us-east-1:filter.tag:Name=mcp-vpc:returnAsymmetricSubnets=true": {
"vpcId": "vpc-0abc123def456789",
"vpcCidrBlock": "10.0.0.0/16",
"subnetGroups": [...]
}
}
Docker asset publishing is handled automatically when a CDK app uses ContainerImage.fromAsset(). CDK Pipelines inserts a PublishAssets stage that builds the Docker image and pushes it to ECR in the pipeline account (or a cross-account ECR if configured). Asset hashing uses the content of the Docker build context — if the Dockerfile and source files are unchanged between commits, CDK produces the same asset hash and the PublishAssets stage skips the push. This means unchanged services in a monorepo incur no ECR push cost on commits that only touch unrelated files.
Consolidated Failure Modes
| # | Failure | Symptom | Root Cause | Fix |
|---|---|---|---|---|
| 1 | CodeStar Connection not AVAILABLE | Pipeline execution fails at Source stage immediately with authorization error | Connection created but OAuth flow not completed — connection is in PENDING state | Authorize the connection in the AWS Console (CodeStar Connections → Update pending connection) before the first pipeline execution; the AVAILABLE state is required and cannot be set via CLI or CloudFormation |
| 2 | Pipeline service role missing iam:PassRole |
Deploy stage fails with AccessDenied: User is not authorized to perform iam:PassRole |
iam:PassRole not scoped correctly, or condition on iam:PassedToService blocks the action |
Scope iam:PassRole to the ECS task execution role ARN with condition iam:PassedToService: ecs-tasks.amazonaws.com; verify the resource ARN matches the exact task execution role used by the task definition |
| 3 | imagedefinitions.json missing from CodeBuild output |
ECS deploy action fails with "No container mapping found in imagedefinitions.json" | Missing or malformed post_build step; container name mismatch; file not included in output artifacts | Add printf '[{"name":"mcp-server","imageUri":"%s"}]' "$ECR_IMAGE" > imagedefinitions.json to post_build; verify container name matches the ECS task definition container name exactly (case-sensitive); add imagedefinitions.json to buildspec artifacts.files |
| 4 | PARALLEL execution mode on same ECS service | Both executions update the same ECS service simultaneously; second wins, first's deployment may be partially applied | PARALLEL mode allows simultaneous pipeline executions that both call UpdateService on the same cluster/service |
Use SUPERSEDED mode (new push cancels in-flight) for nearly all ECS deployments; use QUEUED only when ordered deployment is required; never use PARALLEL for services sharing a deployment target |
| 5 | CodeDeploy agent not on EC2 instance | Instance silently skipped; deployment shows "0 failed" but old code still running | Agent not installed or not running on the instance at deployment time | Bake agent into AMI, or install via SSM RunCommand with AWS-ConfigureAWSPackage document before deploying; add an ASG lifecycle hook to install agent on new instances before they enter the deployment group |
| 6 | appspec.yml file_exists_behavior missing |
BeforeInstall hook fails with "file already exists" error on update deployments | Default file_exists_behavior: DISALLOW refuses to overwrite files from a previous deployment revision |
Add file_exists_behavior: OVERWRITE to each entry in the files section of appspec.yml; OVERWRITE replaces existing files at the destination path |
| 7 | ValidateService hook exits non-zero on slow startup | Deployment rolls back unexpectedly when service is healthy but slow to start | Health check loop too strict with no retry — MCP server takes longer than expected to bind the port | Add retry loop: for i in $(seq 1 10); do curl -sf /health && exit 0; sleep 5; done; exit 1; tune the retry count and sleep duration to the actual p95 startup time of the MCP server |
| 8 | ECS service has deployment controller ECS, trying to attach CodeDeploy | CodeDeploy deployment group creation fails: "cannot apply blue-green to a service with ECS deployment controller" | Deployment controller is immutable after ECS service creation; the existing service was created with the default ECS controller | Recreate the ECS service with deploymentController: { type: CODE_DEPLOY }; update the CloudFormation resource or CDK construct to use CODE_DEPLOY before the first service creation |
| 9 | terminationWaitTimeInMinutes set to 0 |
Blue environment terminated immediately after traffic shift; zero rollback window | Default or explicitly set termination wait time of 0 terminates blue tasks the moment full traffic shift completes | Set terminationWaitTimeInMinutes: 60 for production — keeps blue tasks alive for 60 minutes so you can trigger rollback via aws deploy stop-deployment if post-deploy monitoring reveals issues under real traffic |
| 10 | BeforeAllowTraffic hook times out without calling PutLifecycleEventHookExecutionStatus | Deployment stalls indefinitely at hook stage; does not advance or fail until manual intervention | Lambda function completed (or errored) but did not call codedeploy:PutLifecycleEventHookExecutionStatus before the hook timeout |
Always call PutLifecycleEventHookExecutionStatus in both the success and exception paths of the Lambda — wrap the test logic in try/finally and call it in finally; missing the call in the error path leaves the deployment stalled rather than rolling back |
| 11 | CDK selfMutation loop | Pipeline always shows "pipeline updated" and restarts before deploying any application stage | Non-deterministic synth output — Date.now() in logical IDs, unpinned CDK version changing asset hashes, or uncommitted cdk.context.json |
Pin CDK version in package.json (exact version, not caret range), remove all dynamic values from construct IDs, commit cdk.context.json so context lookups are stable across synth runs |
| 12 | Cross-account cdk bootstrap without --trust-for-lookup |
Synth fails with "Context lookup failed" or deploy fails with "CDKToolkit stack not found in target account" | Target account bootstrapped with --trust only; lookup role not trusted to pipeline account |
Re-run cdk bootstrap --trust <pipeline-account> --trust-for-lookup <pipeline-account> in every target account; --trust-for-lookup is required for VPC ID, AMI ID, and hosted zone lookups during synth |
| 13 | V2 pipeline variable returns empty string | Downstream action receives empty string for #{namespace.variable}; ECS deploy gets malformed image tag |
Variable exported from CodeBuild but not listed in exported-variables array in buildspec, or namespace not declared on the Build action |
Add variable name to both env.exported-variables and env.variables (or set it via shell in commands); verify the action's Namespace property matches the #{namespace.variable} prefix used downstream |
| 14 | Manual Approval SNS topic missing codepipeline.amazonaws.com publish permission |
No approval notification sent but pipeline waits indefinitely; appears as hung pipeline | SNS topic policy does not grant the CodePipeline service principal permission to publish messages | Add Principal: {Service: codepipeline.amazonaws.com} to the SNS topic policy with Action: sns:Publish and a Condition: ArnLike: aws:SourceArn: arn:aws:codepipeline:...:pipeline-name for least privilege |
| 15 | ALB deregistration delay too short for MCP SSE connections | In-flight MCP tool calls dropped during blue-green traffic shift; clients receive connection reset errors | Default 300-second deregistration delay does not account for long-running SSE connections that persist for the lifetime of an MCP session | Increase deregistration_delay.timeout_seconds to 600 or more on the blue target group; also ensure the MCP server's SSE handler sends SIGTERM gracefully and drains in-flight responses before exiting |
Production Checklists
CodePipeline Setup Checklist
- Pipeline type is V2 (required for pipeline variables, path filter triggers, and execution modes beyond SUPERSEDED)
- CodeStar Connection is in
AVAILABLEstate — manually complete the OAuth flow in the AWS Console before first use - Artifact store S3 bucket has versioning enabled and a lifecycle rule expiring objects after 30 days and noncurrent versions after 7 days
- Artifact store uses a KMS CMK; pipeline service role has
kms:GenerateDataKeyandkms:Decrypton the CMK ARN - Pipeline service role has
iam:PassRolescoped to the ECS task execution role ARN with conditioniam:PassedToService: ecs-tasks.amazonaws.com - Source action uses
OutputArtifactFormat: CODEBUILD_CLONE_REFfor SHA-based image tags; CodeBuild project role hascodestar-connections:UseConnection - CodeBuild post_build phase writes
imagedefinitions.jsonwith the correct container name matching the ECS task definition - Execution mode is
SUPERSEDEDfor ECS deploys;PARALLELmode is not used for pipelines deploying to the same ECS service - V2 trigger filters use
push.filePaths.includesfor monorepo path filtering; expected paths verified against actual changed file paths - Build stage
Namespaceproperty set; all exported variables listed in bothenv.exported-variablesand buildspec commands
CodeDeploy EC2 Checklist
- CodeDeploy agent installed and running on all EC2 instances in the deployment group before first deployment
- ASG lifecycle hook or AMI bake ensures new instances have the agent before joining the deployment group
appspec.ymllocated at the bundle root (top level of the deployment artifact ZIP)- Every
filesentry hasfile_exists_behavior: OVERWRITEto handle update deployments - ValidateService hook has a retry loop (10 retries, 5-second sleep) to account for slow MCP server startup
- ValidateService exit code is non-zero on failure (triggers automatic rollback) and zero on success
- Deployment configuration is not
AllAtOncefor production;OneAtATimefor safest rollout - Automatic rollback enabled on
DEPLOYMENT_FAILUREandDEPLOYMENT_STOP_ON_ALARMwith a CloudWatch alarm name specified - BeforeInstall hook sends SIGTERM for graceful drain of MCP SSE connections with a fallback SIGKILL timeout
ECS Blue-Green Checklist
- ECS service created with
deploymentController: { type: CODE_DEPLOY }— cannot be changed after service creation - Two ALB target groups registered: blue (production listener port 443) and green (test listener port 8443)
- Test listener exists on a separate port from the production listener — smoke tests run against the test listener during
BeforeAllowTraffic - CodeDeploy deployment group references both target groups and both listener ARNs
appspec.yaml(not.yml) contains the<TASK_DEFINITION>literal placeholder- Build artifact includes
appspec.yaml,taskdef.json, andimagedefinitions.json terminationWaitTimeInMinutesset to 60 or more for production environments- BeforeAllowTraffic Lambda hook always calls
codedeploy:PutLifecycleEventHookExecutionStatusin both success and exception paths (wrap in try/finally) - Traffic shifting policy is
ECSLinear10PercentEvery1Minutesfor production (notECSAllAtOnce) - ALB deregistration delay on the blue target group set to 600 seconds or more to avoid dropping in-flight MCP SSE connections
CDK Pipelines Checklist
- Import from
aws-cdk-lib/pipelines— notaws-codepipeline selfMutation: trueset on theCodePipelineconstruct- Synth step
primaryOutputDirectorymatches the CDK output directory (defaultcdk.out, or the directory containing it ifcdk synthis run from a subdirectory) - CDK version pinned to an exact version in
package.json(no caret or tilde ranges) to prevent self-mutation from version-bump-induced hash changes - No
Date.now(),Math.random(), or other non-deterministic values in construct logical IDs cdk.context.jsoncommitted to the repository; context lookups stable across synth runs- Cross-account bootstrap run in every target account with both
--trustand--trust-for-lookuppointing to the pipeline account ID - Bootstrap version in target accounts meets the minimum version required by the CDK version in use
- Multi-region deployment uses
pipeline.addWave()with oneaddStage()per region to deploy all regions in parallel dockerEnabledForSynth: trueset if synth step builds Docker images or if app usesContainerImage.fromAsset()