AWS CI/CD · 2026-10-08 · CodePipeline/CodeDeploy arc

AWS CI/CD for MCP Servers: CodePipeline V2, CodeDeploy Blue-Green, and CDK Pipelines

The single most disruptive mistake in a CodePipeline setup is creating a GitHub V2 source action against a CodeStar Connection that is still in PENDING state — the pipeline creates without error, but every execution attempt fails immediately at Source with a cryptic authorization message that looks like an IAM problem, when the real fix is a one-time manual OAuth flow in the AWS Console that turns the connection to AVAILABLE. For MCP server deployments, CodePipeline solves the "push to GitHub and have a tested revision running in ECS within minutes" problem — a V2 pipeline with a CodeStar connection source, a CodeBuild build stage that produces an imagedefinitions.json artifact, and an ECS blue-green deploy action with CodeDeploy gives you atomic traffic shifting and a 60-minute rollback window with zero infrastructure management. The five topics that produce a production-ready CI/CD pipeline for MCP servers — CodePipeline V2 structure and GitHub integration (pipeline type V2, CodeStar Connection, artifact store, service role, execution modes, imagedefinitions.json, monorepo path filters), stage design patterns (action categories, runOrder, V2 pipeline variables, Manual Approval, staging-to-prod flow), CodeDeploy EC2 rolling deployments (CodeDeploy agent, appspec.yml structure, lifecycle hooks, deployment configurations, automatic rollback), ECS blue-green deployments (dual target groups, traffic shifting policies, terminationWaitTimeInMinutes, BeforeAllowTraffic/AfterAllowTraffic hooks, ALB deregistration delay for SSE), and CDK Pipelines self-mutation (selfMutation, synth step, cross-account bootstrap, addStage, waves for parallel multi-region rollout) — each contains sharp edges that produce silent failures, permanently stuck deployments, or infrastructure drift that only surfaces after a failed rollback. The pipeline service role needs iam:PassRole scoped to the ECS task execution role with a condition limiting the pass to ecs-tasks.amazonaws.com — without that condition the role is effectively unrestricted and will fail an AWS Security Hub finding for overly permissive role-passing. ECS blue-green requires the service to have been created with deploymentController: { type: CODE_DEPLOY } — you cannot switch the deployment controller type in-place after service creation, so retrofitting blue-green onto an existing ECS service means destroying and recreating the service. This guide synthesizes all five topics into three structural patterns: pipeline setup and GitHub integration, deployment strategies (rolling EC2, ECS blue-green, approval gates), and CDK Pipelines infrastructure as code with self-mutation.

TL;DR

Pattern 1 — Pipeline Setup and GitHub Integration

CodePipeline V2 and the CodeStar Connection authorization gap

CodePipeline has two pipeline types: V1 (legacy) and V2 (current). Use V2 for all new MCP server pipelines — V2 adds pipeline-level variables that can be declared on one action and consumed in downstream stages, and adds trigger filters including push.filePaths.includes for monorepo path filtering so pushes touching only docs/ do not trigger a full rebuild and deploy. The V1 type will not receive new features and AWS has been quietly guiding customers to V2 since 2023.

The GitHub V2 source action uses CodeStar Connections, which is the correct mechanism for all new GitHub integrations (the older GitHub V1 action used OAuth tokens stored in Secrets Manager and is deprecated). The critical property of a CodeStar Connection that trips nearly every first-time setup: a newly created connection starts in PENDING state. The pipeline creates successfully, the source action saves without error — but every pipeline execution attempt fails immediately at the Source stage because the connection has not been authorized. The fix is a one-time manual step: navigate to the CodeStar Connections console, select the connection, click "Update pending connection," and complete the GitHub OAuth flow. Once completed, the connection reaches AVAILABLE state and all future executions succeed. There is no AWS CLI or CloudFormation mechanism to complete this OAuth step — it must be done by a human in the console before the pipeline is used in production.

# Create a CodeStar Connection for GitHub
aws codestar-connections create-connection \
  --provider-type GitHub \
  --connection-name mcp-server-github \
  --region us-east-1
# Returns: ConnectionArn — status will be PENDING

# After manually authorizing in the console, verify AVAILABLE state:
aws codestar-connections get-connection \
  --connection-arn arn:aws:codestar-connections:us-east-1:123456789012:connection/abc123 \
  --query 'Connection.ConnectionStatus' \
  --output text
# Must return: AVAILABLE (not PENDING)

# CloudFormation: V2 pipeline with GitHub V2 source
# PipelineType: V2 is required for variables and path filters
Resources:
  McpPipeline:
    Type: AWS::CodePipeline::Pipeline
    Properties:
      PipelineType: V2
      RoleArn: !GetAtt PipelineRole.Arn
      ArtifactStore:
        Type: S3
        Location: !Ref ArtifactBucket
        EncryptionKey:
          Type: KMS
          Id: !Ref ArtifactKMSKey
      Triggers:
        - ProviderType: CodeStarSourceConnection
          GitConfiguration:
            SourceActionName: GitHub
            Push:
              - Branches:
                  Includes: [main]
                FilePaths:
                  Includes: [src/**, infrastructure/**]
                  # Unmatched pushes (e.g. docs-only) are silently skipped
      Stages:
        - Name: Source
          Actions:
            - Name: GitHub
              ActionTypeId:
                Category: Source
                Owner: AWS
                Provider: CodeStarSourceConnection
                Version: "1"
              Configuration:
                ConnectionArn: !Ref GitHubConnection
                FullRepositoryId: myorg/mcp-server
                BranchName: main
                OutputArtifactFormat: CODEBUILD_CLONE_REF
                # Full git clone with history for SHA-based image tags
                # Requires CodeBuild role to also have codestar-connections:UseConnection
              OutputArtifacts:
                - Name: SourceArtifact
        - Name: Build
          Actions:
            - Name: BuildAndPush
              ActionTypeId:
                Category: Build
                Owner: AWS
                Provider: CodeBuild
                Version: "1"
              Namespace: BuildVars
              Configuration:
                ProjectName: !Ref CodeBuildProject
              InputArtifacts:
                - Name: SourceArtifact
              OutputArtifacts:
                - Name: BuildArtifact

Artifact store and pipeline service role

The artifact store is an S3 bucket that CodePipeline uses to pass artifacts between stages. Every artifact — source code ZIP, build output, appspec, imagedefinitions.json — passes through this bucket. Three configuration requirements matter for a production MCP server pipeline: versioning must be enabled (CodePipeline requires it and will fail to create without it), encryption should use a KMS CMK so you can audit key usage and rotate the key independently of the bucket, and a lifecycle rule should expire objects after 30 days (noncurrent versions after 7 days) to prevent unbounded storage growth from every push ever made to the branch.

The pipeline service role requires precise permissions. The most common gap is iam:PassRole — when the ECS deploy action registers a new task definition and updates the ECS service, it passes the ECS task execution role to ECS. The pipeline service role must have iam:PassRole on that task execution role ARN. Scope the permission with the condition iam:PassedToService: ecs-tasks.amazonaws.com to prevent the role from being used to pass any role to any service — without this condition, a broadly scoped iam:PassRole is a privilege escalation vector.

# Pipeline service role inline policy
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "CodeStarConnection",
      "Effect": "Allow",
      "Action": "codestar-connections:UseConnection",
      "Resource": "arn:aws:codestar-connections:us-east-1:123456789012:connection/*"
    },
    {
      "Sid": "CodeBuild",
      "Effect": "Allow",
      "Action": ["codebuild:StartBuild", "codebuild:BatchGetBuilds", "codebuild:StopBuild"],
      "Resource": "arn:aws:codebuild:us-east-1:123456789012:project/mcp-server-build"
    },
    {
      "Sid": "ECSDeployRead",
      "Effect": "Allow",
      "Action": [
        "ecs:RegisterTaskDefinition",
        "ecs:DescribeTaskDefinition",
        "ecs:DescribeServices"
      ],
      "Resource": "*"
    },
    {
      "Sid": "ECSDeployWrite",
      "Effect": "Allow",
      "Action": ["ecs:UpdateService"],
      "Resource": "arn:aws:ecs:us-east-1:123456789012:service/mcp-cluster/mcp-server"
    },
    {
      "Sid": "PassRoleToECS",
      "Effect": "Allow",
      "Action": "iam:PassRole",
      "Resource": "arn:aws:iam::123456789012:role/McpTaskExecutionRole",
      "Condition": {
        "StringEquals": { "iam:PassedToService": "ecs-tasks.amazonaws.com" }
      }
    },
    {
      "Sid": "ArtifactStoreKMS",
      "Effect": "Allow",
      "Action": ["kms:GenerateDataKey", "kms:Decrypt"],
      "Resource": "arn:aws:kms:us-east-1:123456789012:key/artifact-key-id"
    },
    {
      "Sid": "ArtifactStoreS3",
      "Effect": "Allow",
      "Action": ["s3:GetObject", "s3:PutObject", "s3:GetObjectVersion"],
      "Resource": "arn:aws:s3:::mcp-pipeline-artifacts/*"
    },
    {
      "Sid": "SNSApproval",
      "Effect": "Allow",
      "Action": "sns:Publish",
      "Resource": "arn:aws:sns:us-east-1:123456789012:mcp-deploy-approvals"
    }
  ]
}

CodeBuild stage and imagedefinitions.json

The build stage is where Docker images are built, tagged with the git commit SHA, pushed to ECR, and the deployment artifact is assembled. The commit SHA as the image tag is the right approach for MCP server images: it makes every image version traceable to a specific commit, prevents the "latest" ambiguity where two environments might run images with the same tag but different code, and simplifies rollback (revert the tag in the task definition).

The imagedefinitions.json file is the bridge between CodeBuild and the ECS deploy action. It maps container names to ECR image URIs and is written in the post_build phase. A missing file, a file in the wrong output artifact, or a file with mismatched container name causes the ECS deploy action to fail with a generic "no container mapping found" error that does not indicate which artifact or container name is wrong — so validate both the container name (it must exactly match the container name in the ECS task definition, case-sensitive) and that the file is included in the CodeBuild output artifact.

# buildspec.yml for MCP server build stage
version: 0.2

env:
  variables:
    REPO_NAME: mcp-server
    AWS_DEFAULT_REGION: us-east-1
  exported-variables:
    - IMAGE_TAG
    - ECR_IMAGE

phases:
  pre_build:
    commands:
      - echo Logging in to Amazon ECR...
      - aws ecr get-login-password --region $AWS_DEFAULT_REGION | docker login --username AWS --password-stdin $AWS_ACCOUNT_ID.dkr.ecr.$AWS_DEFAULT_REGION.amazonaws.com
      - COMMIT_SHA=$(git rev-parse --short HEAD)
      - IMAGE_TAG="${COMMIT_SHA}-${CODEBUILD_BUILD_NUMBER}"
      - ECR_REPO="$AWS_ACCOUNT_ID.dkr.ecr.$AWS_DEFAULT_REGION.amazonaws.com/$REPO_NAME"
      - ECR_IMAGE="$ECR_REPO:$IMAGE_TAG"

  build:
    commands:
      - echo Building Docker image at $(date)
      - docker build -t $ECR_IMAGE -f Dockerfile .
      - docker tag $ECR_IMAGE $ECR_REPO:latest

  post_build:
    commands:
      - echo Pushing Docker image...
      - docker push $ECR_IMAGE
      - docker push $ECR_REPO:latest
      # imagedefinitions.json: container name MUST match ECS task definition container name exactly
      - printf '[{"name":"mcp-server","imageUri":"%s"}]' "$ECR_IMAGE" > imagedefinitions.json
      # For multiple containers in the same task:
      # printf '[{"name":"mcp-server","imageUri":"%s"},{"name":"sidecar","imageUri":"%s"}]' "$ECR_IMAGE" "$SIDECAR_IMAGE" > imagedefinitions.json

artifacts:
  files:
    - imagedefinitions.json
    - appspec.yaml       # required for ECS CodeDeploy blue-green
    - taskdef.json       # required for ECS CodeDeploy blue-green
  discard-paths: yes

For ECS blue-green deployments via CodeDeploy, the build artifact must include three files: imagedefinitions.json (for standard ECS rolling), appspec.yaml (CodeDeploy deployment specification), and taskdef.json (the ECS task definition JSON with a <IMAGE1_NAME> placeholder that CodeDeploy replaces). Use jq in the post_build phase to inject the new image URI into the task definition: jq --arg IMAGE "$ECR_IMAGE" '.containerDefinitions[0].image = $IMAGE' taskdef.template.json > taskdef.json. The appspec.yaml for ECS should contain the <TASK_DEFINITION> placeholder that instructs CodeDeploy to register the taskdef.json from the artifact as a new task revision.

Execution modes and monorepo path filtering

CodePipeline V2 supports three execution modes. SUPERSEDED means a new pipeline execution cancels any in-flight execution of the same pipeline — this is the correct mode for most MCP server deployments because you only care about shipping the latest commit, not every intermediate commit. QUEUED means every commit is deployed in order, with each execution waiting for the previous to complete — useful when each commit represents a database migration or other ordered operation. PARALLEL allows multiple simultaneous executions of the same pipeline — never use this for ECS services where two parallel executions would call UpdateService on the same cluster and service simultaneously; the second call wins but the first's deployment may be partially applied, leaving the service in an inconsistent state.

The V2 trigger filter push.filePaths.includes enables monorepo path filtering: pushes that touch no matching path are silently skipped — no execution is created, no failure is recorded. This is correct behavior (unmatched pushes should not trigger the pipeline) but can be confusing when debugging: if a push does not trigger the pipeline, verify the changed file paths against the filter configuration rather than assuming an IAM or webhook problem.

Pattern 2 — Deployment Strategies: Rolling EC2, ECS Blue-Green, Approval Gates

CodeDeploy EC2 rolling deployments

The foundational requirement for CodeDeploy EC2 deployments is the CodeDeploy agent running on every instance in the deployment group. The agent is a daemon process that polls the CodeDeploy service for deployment instructions and executes lifecycle hooks locally. New Auto Scaling Group instances do not have the agent unless it is baked into the AMI or installed via an ASG lifecycle hook using SSM Run Command with the AWS-ConfigureAWSPackage document. The failure mode is silent: instances without the agent are skipped, the deployment reports zero failures and zero successes for those instances, and the old code revision keeps running. This looks like a successful deployment from the CodePipeline perspective.

# Install CodeDeploy agent via SSM Run Command (for existing instances)
aws ssm send-command \
  --document-name "AWS-ConfigureAWSPackage" \
  --parameters '{"action":["Install"],"name":["AWSCodeDeployAgent"]}' \
  --targets '[{"Key":"tag:Env","Values":["production"]}]' \
  --region us-east-1

# appspec.yml — must be at the bundle root (same directory level as your app files)
version: 0.0
os: linux
files:
  - source: /
    destination: /opt/mcp-server
    file_exists_behavior: OVERWRITE
    # OVERWRITE required — default DISALLOW refuses to overwrite existing files
    # on update deployments and causes BeforeInstall to fail

permissions:
  - object: /opt/mcp-server
    owner: mcp
    group: mcp
    mode: "755"
    type:
      - directory
  - object: /opt/mcp-server
    pattern: "**"
    owner: mcp
    group: mcp
    mode: "644"
    type:
      - file

hooks:
  BeforeInstall:
    - location: scripts/stop_server.sh
      timeout: 60
      runas: root
  AfterInstall:
    - location: scripts/install_deps.sh
      timeout: 120
      runas: mcp
  ApplicationStart:
    - location: scripts/start_server.sh
      timeout: 30
      runas: root
  ValidateService:
    - location: scripts/validate_service.sh
      timeout: 120
      runas: root

The lifecycle hook scripts define the deployment choreography. The BeforeInstall hook should stop the running MCP server gracefully — send SIGTERM first (which triggers graceful drain of open SSE connections) and fall back to SIGKILL after a timeout. The AfterInstall hook installs or updates dependencies (e.g., npm ci --production or pip install -r requirements.txt). ApplicationStart starts the service (typically by enabling and starting a systemd unit). ValidateService is the most important hook: it must verify the new revision is actually serving requests before CodeDeploy considers the deployment successful.

#!/bin/bash
# scripts/stop_server.sh
set -e
systemctl stop mcp-server || true
# Wait for graceful shutdown of SSE connections (up to 30 seconds)
for i in $(seq 1 30); do
  if ! systemctl is-active --quiet mcp-server; then
    echo "Server stopped gracefully"
    exit 0
  fi
  sleep 1
done
# Force stop if still running
systemctl kill -s SIGKILL mcp-server || true

---
#!/bin/bash
# scripts/validate_service.sh
# Non-zero exit triggers automatic rollback
set -e
MAX_RETRIES=10
HEALTH_URL="http://localhost:8080/health"
for i in $(seq 1 $MAX_RETRIES); do
  if curl -sf "$HEALTH_URL" --max-time 5; then
    echo "Health check passed on attempt $i"
    exit 0
  fi
  echo "Attempt $i failed, retrying in 5s..."
  sleep 5
done
echo "Health check failed after $MAX_RETRIES attempts"
exit 1

Deployment configurations control how many instances are updated simultaneously. CodeDeployDefault.OneAtATime keeps all minus one instance healthy throughout the deployment (safest for production — never takes more than one instance offline), CodeDeployDefault.HalfAtATime takes down 50% at a time for a two-round rollout, and CodeDeployDefault.AllAtOnce takes all instances offline simultaneously — never use AllAtOnce in production because a failed deployment leaves zero healthy instances serving traffic. Enable automatic rollback on DEPLOYMENT_FAILURE and DEPLOYMENT_STOP_ON_ALARM with a CloudWatch alarm that fires when the application error rate exceeds a threshold — this gives you automated rollback when the ValidateService hook passes but the new revision develops errors under real traffic.

ECS blue-green deployments with CodeDeploy

ECS blue-green is the deployment strategy for production MCP servers on ECS: the new task revision starts in a parallel "green" environment behind an ALB test listener, traffic shifts from the old "blue" environment to green only after health checks pass, and the blue environment remains live for a configurable window in case rollback is needed. The constraint that catches teams by surprise is that the deployment controller type on an ECS service is immutable — a service created with deploymentController: { type: ECS }` (the default) cannot be changed to `CODE_DEPLOY without recreating the service. Plan the deployment controller at service creation time.

# ECS service requiring CODE_DEPLOY deployment controller for blue-green
aws ecs create-service \
  --cluster mcp-cluster \
  --service-name mcp-server \
  --task-definition mcp-server:1 \
  --desired-count 2 \
  --launch-type FARGATE \
  --deployment-controller type=CODE_DEPLOY \
  --network-configuration "awsvpcConfiguration={subnets=[subnet-abc,subnet-def],securityGroups=[sg-123],assignPublicIp=DISABLED}" \
  --load-balancers '[
    {
      "targetGroupArn": "arn:aws:elasticloadbalancing:us-east-1:123456789012:targetgroup/mcp-blue/abc123",
      "containerName": "mcp-server",
      "containerPort": 8080
    }
  ]'

# appspec.yaml for ECS blue-green deployment
# Note: .yaml extension required (not .yml) for ECS CodeDeploy
version: 0.0
Resources:
  - TargetService:
      Type: AWS::ECS::Service
      Properties:
        TaskDefinition: 
        #  is a literal placeholder — CodeDeploy replaces it
        # with the ARN of the task revision registered from taskdef.json
        LoadBalancerInfo:
          ContainerName: mcp-server
          ContainerPort: 8080
        PlatformVersion: LATEST
Hooks:
  - BeforeAllowTraffic: McpSmokeTestFunction
    # Lambda ARN or function name — called after green passes health checks
    # but before any production traffic shifts to green
    # Must call codedeploy:PutLifecycleEventHookExecutionStatus
  - AfterAllowTraffic: McpPostDeployFunction

# CodeDeploy deployment group for ECS blue-green
aws deploy create-deployment-group \
  --application-name mcp-server-app \
  --deployment-group-name mcp-server-prod \
  --service-role-arn arn:aws:iam::123456789012:role/CodeDeployECSRole \
  --deployment-config-name CodeDeployDefault.ECSLinear10PercentEvery1Minutes \
  --ecs-services clusterName=mcp-cluster,serviceName=mcp-server \
  --load-balancer-info '{
    "targetGroupPairInfoList": [{
      "targetGroups": [
        {"name": "mcp-blue"},
        {"name": "mcp-green"}
      ],
      "prodTrafficRoute": {"listenerArns": ["arn:aws:elasticloadbalancing:...:listener/app/mcp-alb/.../prod"]},
      "testTrafficRoute": {"listenerArns": ["arn:aws:elasticloadbalancing:...:listener/app/mcp-alb/.../test"]}
    }]
  }' \
  --blue-green-deployment-configuration '{
    "terminateBlueInstancesOnDeploymentSuccess": {
      "action": "TERMINATE",
      "terminationWaitTimeInMinutes": 60
    },
    "deploymentReadyOption": {
      "actionOnTimeout": "CONTINUE_DEPLOYMENT",
      "waitTimeInMinutes": 0
    }
  }'

Traffic shifting policies determine how quickly production traffic moves from blue to green. ECSAllAtOnce shifts all traffic immediately — use only for staging environments. ECSLinear10PercentEvery1Minutes shifts 10% per minute for a 10-minute gradual shift while monitoring error rates — the right choice for production MCP servers where client reconnection on SSE drop is noticeable. ECSCanary10Percent5Minutes sends 10% to green for 5 minutes then shifts the remaining 90% — good for quick validation before full shift. ECSCanary10Percent15Minutes holds 10% longer for deeper canary observation when the MCP server calls downstream services with latency that only appears at sustained load.

The terminationWaitTimeInMinutes: 60 setting is critical for MCP servers. After the full traffic shift completes, blue tasks remain alive and registered in the blue target group for 60 minutes. During this window you can trigger a rollback via aws deploy stop-deployment without re-running the full deployment pipeline. After the window closes, blue tasks are terminated and rollback requires a new full deployment from source. Set the ALB deregistration delay on the blue target group to at least 600 seconds (10 minutes) to avoid dropping long-running MCP SSE connections during the initial deregistration phase at traffic shift start.

# BeforeAllowTraffic Lambda hook
# This Lambda runs after green passes health checks, before any traffic shifts
import boto3
import json

codedeploy = boto3.client('codedeploy')

def lambda_handler(event, context):
    deployment_id = event['DeploymentId']
    lifecycle_event_hook_execution_id = event['LifecycleEventHookExecutionId']

    try:
        # Run smoke tests against the test listener (green environment)
        # The test listener routes to green; production listener still routes to blue
        test_alb_endpoint = "https://mcp-test.internal:8443"
        run_smoke_tests(test_alb_endpoint)

        status = 'Succeeded'
    except Exception as e:
        print(f"Smoke test failed: {e}")
        status = 'Failed'
    finally:
        # CRITICAL: always call PutLifecycleEventHookExecutionStatus
        # Failing to call this causes the deployment to stall indefinitely
        codedeploy.put_lifecycle_event_hook_execution_status(
            deploymentId=deployment_id,
            lifecycleEventHookExecutionId=lifecycle_event_hook_execution_id,
            status=status
        )

def run_smoke_tests(endpoint):
    import urllib.request
    req = urllib.request.urlopen(f"{endpoint}/health", timeout=10)
    if req.status != 200:
        raise Exception(f"Health check returned {req.status}")
    # Test that the MCP server responds to tool list requests
    req2 = urllib.request.urlopen(f"{endpoint}/mcp/tools/list", timeout=10)
    if req2.status != 200:
        raise Exception(f"Tool list returned {req2.status}")

Stage design and approval gates

A production stage design for MCP servers typically follows: Source → Build → DeployStaging → TestAndApprove → DeployProd. The TestAndApprove stage uses runOrder to run integration tests in parallel with a Manual Approval action — both actions have runOrder: 1, and the stage does not advance until both succeed. This means an approver cannot accidentally approve a deployment that has failing integration tests (the stage waits for both to succeed simultaneously, not either independently).

# TestAndApprove stage with parallel integration tests and manual approval
- Name: TestAndApprove
  Actions:
    - Name: IntegrationTests
      RunOrder: 1
      ActionTypeId:
        Category: Test
        Owner: AWS
        Provider: CodeBuild
        Version: "1"
      Configuration:
        ProjectName: mcp-integration-tests
      InputArtifacts:
        - Name: BuildArtifact

    - Name: ManualApproval
      RunOrder: 1
      ActionTypeId:
        Category: Approval
        Owner: AWS
        Provider: Manual
        Version: "1"
      Configuration:
        NotificationArn: arn:aws:sns:us-east-1:123456789012:mcp-deploy-approvals
        CustomData: "Deploy #{BuildVars.IMAGE_TAG} to production. Integration test results at #{BuildVars.TEST_REPORT_URL}"
        ExternalEntityLink: "https://github.com/myorg/mcp-server/commit/#{BuildVars.COMMIT_SHA}"
        # #{BuildVars.IMAGE_TAG} references V2 pipeline variable exported from Build stage
        # SNS topic policy must allow codepipeline.amazonaws.com to publish

- Name: DeployProd
  Actions:
    - Name: BlueGreenDeploy
      RunOrder: 1
      ActionTypeId:
        Category: Deploy
        Owner: AWS
        Provider: CodeDeployToECS
        Version: "1"
      Configuration:
        ApplicationName: mcp-server-app
        DeploymentGroupName: mcp-server-prod
        TaskDefinitionTemplateArtifact: BuildArtifact
        TaskDefinitionTemplatePath: taskdef.json
        AppSpecTemplateArtifact: BuildArtifact
        AppSpecTemplatePath: appspec.yaml
        Image1ArtifactName: BuildArtifact
        Image1ContainerName: IMAGE1_NAME
      InputArtifacts:
        - Name: BuildArtifact

The Manual Approval action's SNS topic must have a topic policy granting codepipeline.amazonaws.com permission to publish — without this, the approval notification is silently dropped but the pipeline still waits indefinitely for manual approval, which looks like a hung pipeline rather than a missing notification. Approval tokens expire after 7 days — if no action is taken within 7 days, the execution fails. For teams using Slack, AWS Chatbot integration can deliver approval requests to a Slack channel with one-click approve/reject buttons.

V2 pipeline variables declared via the Namespace property on an action and exported from CodeBuild via exported-variables in the buildspec are consumed downstream as #{namespace.variable}. A common failure is exporting a variable name in exported-variables that does not appear in the buildspec env.variables or env.exported-variables section — the variable resolves to an empty string downstream with no warning, and the deploy action receives a malformed image tag.

Pattern 3 — Infrastructure as Code: CDK Pipelines Self-Mutation

CDK Pipelines construct and selfMutation

CDK Pipelines is an L3 construct for defining CI/CD pipelines that deploy CDK applications. Import it from aws-cdk-lib/pipelines — not from aws-codepipeline. These are entirely separate APIs: aws-cdk-lib/pipelines is the high-level CDK Pipelines construct that manages CodePipeline, CodeBuild, and deployment stages as a single abstraction; aws-codepipeline is the low-level L2 construct for building raw CodePipeline resources. Using the wrong import produces confusing type errors that do not indicate the root cause.

// CDK Pipelines pipeline definition
// File: infrastructure/lib/pipeline-stack.ts
import { Stack, StackProps } from 'aws-cdk-lib';
import { Construct } from 'constructs';
import { CodePipeline, CodePipelineSource, CodeBuildStep, ManualApprovalStep } from 'aws-cdk-lib/pipelines';
// ^^ import from aws-cdk-lib/pipelines — NOT aws-codepipeline
import * as codebuild from 'aws-cdk-lib/aws-codebuild';
import { McpServerStage } from './mcp-server-stage';

export class PipelineStack extends Stack {
  constructor(scope: Construct, id: string, props?: StackProps) {
    super(scope, id, props);

    const pipeline = new CodePipeline(this, 'McpPipeline', {
      pipelineName: 'mcp-server-pipeline',
      selfMutation: true,
      // selfMutation: true adds UpdatePipeline as the first stage
      // after every push, pipeline updates its own CloudFormation stack
      // before deploying any application stage — add stages in TypeScript,
      // push, and the pipeline installs itself without manual cfn update

      crossAccountKeys: true,
      // required for cross-account deployments — creates KMS key
      // shared with target accounts for artifact encryption

      dockerEnabledForSynth: true,
      dockerEnabledForSelfMutation: true,
      // required if synth step builds Docker images or if app uses
      // ContainerImage.fromAsset() — enables Docker daemon in CodeBuild

      synth: new CodeBuildStep('Synth', {
        input: CodePipelineSource.connection('myorg/mcp-server', 'main', {
          connectionArn: 'arn:aws:codestar-connections:us-east-1:123456789012:connection/abc123',
          triggerOnPush: true,
        }),
        installCommands: [
          'npm install -g aws-cdk@2.100.0',
          // Pin CDK version to prevent self-mutation from unpinned package hashes
        ],
        commands: [
          'cd infrastructure',
          'npm ci',
          'npm run build',
          'npx cdk synth',
        ],
        primaryOutputDirectory: 'infrastructure/cdk.out',
        // primaryOutputDirectory must match the CDK output directory
        // mismatch causes pipeline to fail looking for CloudFormation templates
        buildEnvironment: {
          buildImage: codebuild.LinuxBuildImage.STANDARD_7_0,
          privileged: true,  // required for dockerEnabledForSynth
        },
      }),
    });

    // Staging deployment
    const stagingStage = pipeline.addStage(
      new McpServerStage(this, 'Staging', {
        env: { account: '111111111111', region: 'us-east-1' },
      })
    );

    // Post-stage integration tests using CloudFormation outputs
    stagingStage.addPost(
      new CodeBuildStep('IntegrationTests', {
        commands: [
          'npm ci',
          'npx jest --testPathPattern=integration --testTimeout=30000',
        ],
        envFromCfnOutputs: {
          // ALB_DNS_NAME comes from a CfnOutput in McpServerStack
          ALB_DNS_NAME: stagingStage.stackOutputs.albDnsName,
        },
        // envFromCfnOutputs maps CloudFormation stack outputs to env vars
        // in this CodeBuild step — no secondary DynamoDB lookup required
      })
    );

    // Production gate
    stagingStage.addPost(new ManualApprovalStep('ApproveProduction'));

    // Production wave — multi-region parallel deployment
    const prodWave = pipeline.addWave('Production');
    for (const region of ['us-east-1', 'eu-west-1', 'ap-southeast-1']) {
      prodWave.addStage(
        new McpServerStage(this, `Prod-${region}`, {
          env: { account: '222222222222', region },
        })
      );
    }
    // All three region stages deploy in parallel within the Production wave
  }
}

Cross-account bootstrap and context lookups

CDK Pipelines cross-account deployment requires the CDKToolkit bootstrap stack to be installed in every target account. The bootstrap stack creates a deploy role, a CloudFormation execution role, a lookup role, and an asset bucket — all with trust relationships back to the pipeline account. Run the bootstrap command in each target account with the correct trust flags:

# Bootstrap target account (run once per account-region combination)
# Must be run while authenticated to the TARGET account (not the pipeline account)
cdk bootstrap \
  aws://222222222222/us-east-1 \
  --trust 111111111111 \
  --trust-for-lookup 111111111111 \
  --cloudformation-execution-policies arn:aws:iam::aws:policy/AdministratorAccess

# --trust: allows pipeline account to assume the deploy role in this account
# --trust-for-lookup: allows pipeline account to assume the lookup role
#   for VPC ID, AMI ID, and other context lookups during cdk synth
#   WITHOUT --trust-for-lookup, synth fails with:
#   "Cannot use 'Vpc.fromLookup()': Resolution failed: Context lookup failed"

# Also bootstrap the pipeline account itself (for asset publishing)
cdk bootstrap \
  aws://111111111111/us-east-1 \
  --cloudformation-execution-policies arn:aws:iam::aws:policy/AdministratorAccess

# Verify bootstrap version compatibility — CDK requires a minimum qualifier version
# If bootstrap version is too old, deployment fails with:
# "CDKToolkit stack version 6 is below the minimum required version 14"
aws cloudformation describe-stacks \
  --stack-name CDKToolkit \
  --region us-east-1 \
  --profile target-account \
  --query 'Stacks[0].Parameters[?ParameterKey==`BootstrapVersion`].ParameterValue' \
  --output text

Avoiding the self-mutation loop

The self-mutation feature is powerful but has a failure mode that blocks all deployments: if the CDK synth output is non-deterministic, the pipeline perpetually detects that the pipeline stack needs updating, runs UpdatePipeline, and then restarts — never reaching the application deployment stages. The three common causes of non-deterministic synth are: using Date.now() or Math.random() in logical IDs or construct names (generates different values each synth), unpinned CDK versions in package.json (a patch release can change asset hashing behavior, producing different asset keys per synth), and uncommitted context lookups in cdk.context.json (VPC ID lookups performed during synth write to cdk.context.json and subsequent synths read from cache — but if the cache file is not committed, each synth performs a fresh lookup that may resolve to a slightly different context object).

// WRONG — non-deterministic logical ID causes self-mutation loop
export class McpServerStack extends Stack {
  constructor(scope: Construct, id: string, props?: StackProps) {
    super(scope, id, props);

    // NEVER use Date.now() or Math.random() in construct IDs
    new CfnOutput(this, `Output-${Date.now()}`, { value: 'foo' });
    //                     ^^^^^^^^^^^^ different each synth = perpetual loop
  }
}

// CORRECT — deterministic construct IDs
export class McpServerStack extends Stack {
  readonly albDnsName: CfnOutput;

  constructor(scope: Construct, id: string, props?: StackProps) {
    super(scope, id, props);

    const alb = new elbv2.ApplicationLoadBalancer(this, 'McpAlb', {
      vpc,
      internetFacing: true,
    });

    // Fixed logical ID — identical across every synth
    this.albDnsName = new CfnOutput(this, 'AlbDnsName', {
      value: alb.loadBalancerDnsName,
      exportName: `${this.stackName}-AlbDnsName`,
    });
  }
}

// package.json — pin CDK version to prevent hash drift
{
  "dependencies": {
    "aws-cdk-lib": "2.100.0",  // exact version, not ^2.100.0
    "constructs": "10.3.0"     // exact version
  },
  "devDependencies": {
    "aws-cdk": "2.100.0"       // exact version
  }
}

# cdk.context.json — commit this file to the repository
# Context lookups (VPC IDs, AMI IDs, availability zones) are cached here
# Without committed context, each synth performs fresh lookups that may differ
{
  "availability-zones:account=222222222222:region=us-east-1": [
    "us-east-1a",
    "us-east-1b",
    "us-east-1c"
  ],
  "vpc-provider:account=222222222222:region=us-east-1:filter.tag:Name=mcp-vpc:returnAsymmetricSubnets=true": {
    "vpcId": "vpc-0abc123def456789",
    "vpcCidrBlock": "10.0.0.0/16",
    "subnetGroups": [...]
  }
}

Docker asset publishing is handled automatically when a CDK app uses ContainerImage.fromAsset(). CDK Pipelines inserts a PublishAssets stage that builds the Docker image and pushes it to ECR in the pipeline account (or a cross-account ECR if configured). Asset hashing uses the content of the Docker build context — if the Dockerfile and source files are unchanged between commits, CDK produces the same asset hash and the PublishAssets stage skips the push. This means unchanged services in a monorepo incur no ECR push cost on commits that only touch unrelated files.

Consolidated Failure Modes

#FailureSymptomRoot CauseFix
1 CodeStar Connection not AVAILABLE Pipeline execution fails at Source stage immediately with authorization error Connection created but OAuth flow not completed — connection is in PENDING state Authorize the connection in the AWS Console (CodeStar Connections → Update pending connection) before the first pipeline execution; the AVAILABLE state is required and cannot be set via CLI or CloudFormation
2 Pipeline service role missing iam:PassRole Deploy stage fails with AccessDenied: User is not authorized to perform iam:PassRole iam:PassRole not scoped correctly, or condition on iam:PassedToService blocks the action Scope iam:PassRole to the ECS task execution role ARN with condition iam:PassedToService: ecs-tasks.amazonaws.com; verify the resource ARN matches the exact task execution role used by the task definition
3 imagedefinitions.json missing from CodeBuild output ECS deploy action fails with "No container mapping found in imagedefinitions.json" Missing or malformed post_build step; container name mismatch; file not included in output artifacts Add printf '[{"name":"mcp-server","imageUri":"%s"}]' "$ECR_IMAGE" > imagedefinitions.json to post_build; verify container name matches the ECS task definition container name exactly (case-sensitive); add imagedefinitions.json to buildspec artifacts.files
4 PARALLEL execution mode on same ECS service Both executions update the same ECS service simultaneously; second wins, first's deployment may be partially applied PARALLEL mode allows simultaneous pipeline executions that both call UpdateService on the same cluster/service Use SUPERSEDED mode (new push cancels in-flight) for nearly all ECS deployments; use QUEUED only when ordered deployment is required; never use PARALLEL for services sharing a deployment target
5 CodeDeploy agent not on EC2 instance Instance silently skipped; deployment shows "0 failed" but old code still running Agent not installed or not running on the instance at deployment time Bake agent into AMI, or install via SSM RunCommand with AWS-ConfigureAWSPackage document before deploying; add an ASG lifecycle hook to install agent on new instances before they enter the deployment group
6 appspec.yml file_exists_behavior missing BeforeInstall hook fails with "file already exists" error on update deployments Default file_exists_behavior: DISALLOW refuses to overwrite files from a previous deployment revision Add file_exists_behavior: OVERWRITE to each entry in the files section of appspec.yml; OVERWRITE replaces existing files at the destination path
7 ValidateService hook exits non-zero on slow startup Deployment rolls back unexpectedly when service is healthy but slow to start Health check loop too strict with no retry — MCP server takes longer than expected to bind the port Add retry loop: for i in $(seq 1 10); do curl -sf /health && exit 0; sleep 5; done; exit 1; tune the retry count and sleep duration to the actual p95 startup time of the MCP server
8 ECS service has deployment controller ECS, trying to attach CodeDeploy CodeDeploy deployment group creation fails: "cannot apply blue-green to a service with ECS deployment controller" Deployment controller is immutable after ECS service creation; the existing service was created with the default ECS controller Recreate the ECS service with deploymentController: { type: CODE_DEPLOY }; update the CloudFormation resource or CDK construct to use CODE_DEPLOY before the first service creation
9 terminationWaitTimeInMinutes set to 0 Blue environment terminated immediately after traffic shift; zero rollback window Default or explicitly set termination wait time of 0 terminates blue tasks the moment full traffic shift completes Set terminationWaitTimeInMinutes: 60 for production — keeps blue tasks alive for 60 minutes so you can trigger rollback via aws deploy stop-deployment if post-deploy monitoring reveals issues under real traffic
10 BeforeAllowTraffic hook times out without calling PutLifecycleEventHookExecutionStatus Deployment stalls indefinitely at hook stage; does not advance or fail until manual intervention Lambda function completed (or errored) but did not call codedeploy:PutLifecycleEventHookExecutionStatus before the hook timeout Always call PutLifecycleEventHookExecutionStatus in both the success and exception paths of the Lambda — wrap the test logic in try/finally and call it in finally; missing the call in the error path leaves the deployment stalled rather than rolling back
11 CDK selfMutation loop Pipeline always shows "pipeline updated" and restarts before deploying any application stage Non-deterministic synth output — Date.now() in logical IDs, unpinned CDK version changing asset hashes, or uncommitted cdk.context.json Pin CDK version in package.json (exact version, not caret range), remove all dynamic values from construct IDs, commit cdk.context.json so context lookups are stable across synth runs
12 Cross-account cdk bootstrap without --trust-for-lookup Synth fails with "Context lookup failed" or deploy fails with "CDKToolkit stack not found in target account" Target account bootstrapped with --trust only; lookup role not trusted to pipeline account Re-run cdk bootstrap --trust <pipeline-account> --trust-for-lookup <pipeline-account> in every target account; --trust-for-lookup is required for VPC ID, AMI ID, and hosted zone lookups during synth
13 V2 pipeline variable returns empty string Downstream action receives empty string for #{namespace.variable}; ECS deploy gets malformed image tag Variable exported from CodeBuild but not listed in exported-variables array in buildspec, or namespace not declared on the Build action Add variable name to both env.exported-variables and env.variables (or set it via shell in commands); verify the action's Namespace property matches the #{namespace.variable} prefix used downstream
14 Manual Approval SNS topic missing codepipeline.amazonaws.com publish permission No approval notification sent but pipeline waits indefinitely; appears as hung pipeline SNS topic policy does not grant the CodePipeline service principal permission to publish messages Add Principal: {Service: codepipeline.amazonaws.com} to the SNS topic policy with Action: sns:Publish and a Condition: ArnLike: aws:SourceArn: arn:aws:codepipeline:...:pipeline-name for least privilege
15 ALB deregistration delay too short for MCP SSE connections In-flight MCP tool calls dropped during blue-green traffic shift; clients receive connection reset errors Default 300-second deregistration delay does not account for long-running SSE connections that persist for the lifetime of an MCP session Increase deregistration_delay.timeout_seconds to 600 or more on the blue target group; also ensure the MCP server's SSE handler sends SIGTERM gracefully and drains in-flight responses before exiting

Production Checklists

CodePipeline Setup Checklist

CodeDeploy EC2 Checklist

ECS Blue-Green Checklist

CDK Pipelines Checklist