Guide · AWS CodeDeploy · ECS Blue-Green Deployments

CodeDeploy Blue-Green ECS for MCP Servers — Traffic Shifting, Rollback, Test Listener

CodeDeploy blue-green deployments for ECS give you the ability to deploy a new version of your MCP server alongside the existing version, run automated tests against the new version on a test listener port, and then shift traffic — with a configurable rollback window during which you can revert with a single API call. Unlike ECS rolling updates (which replace tasks in-place on the same target group), blue-green creates a second target group ("green"), registers the new task definition's tasks with it, routes production traffic from the ALB to the green target group, and terminates the original ("blue") tasks only after you specify. This eliminates the brief period during rolling updates where both old and new code handle traffic simultaneously — a critical property for MCP servers that have stateful session handling or in-flight long-running tool calls. Critical decisions: choosing a traffic shifting policy (Linear, Canary, or AllAtOnce), configuring the deployment termination wait time (the rollback window), and wiring a CloudWatch alarm to trigger automatic rollback if your MCP server error rate spikes after the shift.

TL;DR

Blue-green ECS deployments require an ALB with two target groups (blue and green), an ECS service configured for CODE_DEPLOY deployment controller, a CodeDeploy application/deployment group of compute type ECS, and an appspec.yaml that references the new task definition ARN. Use CodeDeployDefault.ECSLinear10PercentEvery1Minutes for canary testing; use CodeDeployDefault.ECSAllAtOnce in staging. Configure a deployment termination wait time of 30-60 minutes to give yourself a rollback window after full traffic shift. Wire a CloudWatch alarm on your MCP server's 5xx error rate to trigger automatic rollback. See CodePipeline integration for triggering ECS blue-green deploys from CI and ECS service configuration for task definition patterns.

Prerequisites: dual target groups and ECS service setup

Blue-green deployments require infrastructure that ECS rolling deployments do not: a second target group on the ALB, a test listener (optional but recommended), and an ECS service explicitly configured for the CodeDeploy deployment controller. You cannot switch an existing ECS service from the ECS deployment controller to CODE_DEPLOY in-place — you must create a new service or delete and recreate.

# Create two target groups — one blue (active), one green (standby)
# CodeDeploy swaps which one receives production ALB traffic
aws elbv2 create-target-group \
  --name mcp-server-blue-tg \
  --protocol HTTP \
  --port 3000 \
  --vpc-id vpc-abc123 \
  --target-type ip \
  --health-check-path /health \
  --health-check-interval-seconds 10 \
  --healthy-threshold-count 2 \
  --unhealthy-threshold-count 3

aws elbv2 create-target-group \
  --name mcp-server-green-tg \
  --protocol HTTP \
  --port 3000 \
  --vpc-id vpc-abc123 \
  --target-type ip \
  --health-check-path /health \
  --health-check-interval-seconds 10 \
  --healthy-threshold-count 2 \
  --unhealthy-threshold-count 3

# Create production listener (port 443) pointing to blue initially
aws elbv2 create-listener \
  --load-balancer-arn $ALB_ARN \
  --protocol HTTPS --port 443 \
  --certificates CertificateArn=$CERT_ARN \
  --default-actions Type=forward,TargetGroupArn=$BLUE_TG_ARN

# Create test listener (port 8443) for green group validation
aws elbv2 create-listener \
  --load-balancer-arn $ALB_ARN \
  --protocol HTTPS --port 8443 \
  --certificates CertificateArn=$CERT_ARN \
  --default-actions Type=forward,TargetGroupArn=$GREEN_TG_ARN
# ECS service with CODE_DEPLOY controller — blue target group is the initial active
aws ecs create-service \
  --cluster mcp-cluster \
  --service-name mcp-server-service \
  --task-definition mcp-server:1 \
  --desired-count 3 \
  --deployment-controller '{"type":"CODE_DEPLOY"}' \
  --network-configuration '{
    "awsvpcConfiguration": {
      "subnets": ["subnet-aaa", "subnet-bbb"],
      "securityGroups": ["sg-mcp-server"],
      "assignPublicIp": "DISABLED"
    }
  }' \
  --load-balancers '[
    {
      "targetGroupArn": "arn:aws:elasticloadbalancing:...:targetgroup/mcp-server-blue-tg/abc",
      "containerName": "mcp-server",
      "containerPort": 3000
    }
  ]'

Once the ECS service is created with CODE_DEPLOY controller, you cannot update the service's task definition using aws ecs update-service --task-definition — that call is only valid for ECS-controller services. All task definition updates must go through CodeDeploy deployments via aws deploy create-deployment.

appspec.yaml for ECS blue-green

The appspec.yaml for ECS blue-green deployments is simpler than the EC2 appspec — there are no file copies or shell scripts, only a task definition reference and optional Lambda hook functions. The task definition is specified as an ARN or as a taskdef.json file in the build artifact.

# appspec.yaml — root of CodePipeline build artifact (alongside imagedefinitions.json)
version: 0.0
Resources:
  - TargetService:
      Type: AWS::ECS::Service
      Properties:
        TaskDefinition: "arn:aws:ecs:us-east-1:123456789012:task-definition/mcp-server:42"
        LoadBalancerInfo:
          ContainerName: "mcp-server"
          ContainerPort: 3000
        PlatformVersion: "LATEST"
Hooks:
  - BeforeAllowTraffic: "mcp-server-pre-traffic-hook"
  - AfterAllowTraffic: "mcp-server-post-traffic-hook"

When used in a CodePipeline, the task definition ARN is typically dynamic — the CodeBuild stage registers a new task definition and writes its ARN to the artifact. There are two patterns for this:

# Pattern 1: Write taskdef.json to the artifact, let CodeDeploy register it
# buildspec.yml post_build:
post_build:
  commands:
    - IMAGE_URI="$ECR_REGISTRY/$ECR_REPO:$CODEBUILD_RESOLVED_SOURCE_VERSION"
    # Inject new image URI into the task definition template
    - >
      jq --arg IMAGE "$IMAGE_URI" '.containerDefinitions[0].image = $IMAGE' \
         taskdef-template.json > taskdef.json
    # Write imagedefinitions.json for ECS rolling deploys (if used)
    - echo "[{\"name\":\"mcp-server\",\"imageUri\":\"$IMAGE_URI\"}]" > imagedefinitions.json
    # appspec.yaml references  placeholder; CodeDeploy resolves it
    - cp appspec-template.yaml appspec.yaml

# appspec-template.yaml with placeholder:
# TaskDefinition: 
# CodeDeploy registers taskdef.json and substitutes its ARN for 

artifacts:
  files:
    - appspec.yaml
    - taskdef.json
    - imagedefinitions.json

The <TASK_DEFINITION> placeholder in the appspec causes CodeDeploy to register the taskdef.json file as a new task definition revision and substitute the resulting ARN. This is the standard pattern for CodePipeline-driven ECS blue-green deployments — it keeps the task definition registration inside the pipeline rather than requiring a separate script.

Traffic shifting policies

The traffic shifting policy controls how CodeDeploy moves production traffic from the blue target group to the green target group. All-at-once shifts 100% of traffic in a single step; linear and canary policies shift traffic gradually with health validation between steps.

PolicyTraffic shift behaviorWhen to use
ECSAllAtOnce100% shift in one step immediatelyStaging environments, internal tools with no SLA, fast deploys where you trust the new build
ECSLinear10PercentEvery1Minutes+10% every minute; full shift in 10 minutesProduction MCP servers with active traffic; errors show up in CloudWatch within 1-2 minutes of exposure
ECSLinear10PercentEvery3Minutes+10% every 3 minutes; full shift in 30 minutesHigh-traffic MCP servers where you want more observation time at each step
ECSCanary10Percent5Minutes10% for 5 minutes, then 100%Quick canary check — validates the new version isn't catastrophically broken before full shift
ECSCanary10Percent15Minutes10% for 15 minutes, then 100%Deeper canary validation; good for MCP servers where errors manifest under sustained load
# Create CodeDeploy application for ECS compute platform
aws deploy create-application \
  --application-name mcp-server-ecs-app \
  --compute-platform ECS

# Create deployment group with traffic shifting and alarm rollback
aws deploy create-deployment-group \
  --application-name mcp-server-ecs-app \
  --deployment-group-name mcp-server-production \
  --deployment-config-name CodeDeployDefault.ECSCanary10Percent5Minutes \
  --service-role-arn arn:aws:iam::123456789012:role/CodeDeployECSServiceRole \
  --ecs-services '[{
    "serviceName": "mcp-server-service",
    "clusterName": "mcp-cluster"
  }]' \
  --load-balancer-info '{
    "targetGroupPairInfoList": [{
      "targetGroups": [
        {"name": "mcp-server-blue-tg"},
        {"name": "mcp-server-green-tg"}
      ],
      "prodTrafficRoute": {
        "listenerArns": ["arn:aws:elasticloadbalancing:...:listener/app/mcp-alb/prod-listener-arn"]
      },
      "testTrafficRoute": {
        "listenerArns": ["arn:aws:elasticloadbalancing:...:listener/app/mcp-alb/test-listener-arn"]
      }
    }]
  }' \
  --auto-rollback-configuration '{
    "enabled": true,
    "events": ["DEPLOYMENT_FAILURE", "DEPLOYMENT_STOP_ON_ALARM"]
  }' \
  --alarm-configuration '{
    "enabled": true,
    "alarms": [
      {"name": "mcp-server-5xx-rate-high"},
      {"name": "mcp-server-latency-p99-high"}
    ]
  }' \
  --blue-green-deployment-configuration '{
    "terminateBlueInstancesOnDeploymentSuccess": {
      "action": "TERMINATE",
      "terminationWaitTimeInMinutes": 60
    },
    "deploymentReadyOption": {
      "actionOnTimeout": "CONTINUE_DEPLOYMENT",
      "waitTimeInMinutes": 0
    }
  }'

The terminationWaitTimeInMinutes: 60 setting means CodeDeploy keeps the blue task set running for 60 minutes after full traffic shift before terminating it. During this window, you can trigger an immediate rollback via the console or aws deploy stop-deployment --auto-rollback-enabled. After the window closes, the old tasks are terminated and rollback requires a new full deployment of the previous image — not a simple swap.

BeforeAllowTraffic and AfterAllowTraffic Lambda hooks

ECS blue-green deployments support two Lambda lifecycle hook points. BeforeAllowTraffic runs after the green task set is up and passing ALB health checks but before any production traffic is shifted. AfterAllowTraffic runs after 100% of traffic has shifted to green. Both hooks must call codedeploy:PutLifecycleEventHookExecutionStatus with Succeeded or Failed to unblock the deployment.

# Lambda hook for BeforeAllowTraffic — smoke test via test listener
import boto3, json, urllib.request

codedeploy = boto3.client('codedeploy')

def handler(event, context):
    deployment_id = event['DeploymentId']
    hook_execution_id = event['LifecycleEventHookExecutionId']

    try:
        # Test listener routes to green target group on port 8443
        req = urllib.request.Request('https://mcp-alb.us-east-1.elb.amazonaws.com:8443/health',
                                     headers={'Host': 'mcp.alivemcp.com'})
        response = urllib.request.urlopen(req, timeout=5)
        body = json.loads(response.read())

        if response.status == 200 and body.get('status') == 'ok':
            status = 'Succeeded'
        else:
            status = 'Failed'
    except Exception as e:
        print(f"Health check error: {e}")
        status = 'Failed'

    codedeploy.put_lifecycle_event_hook_execution_status(
        deploymentId=deployment_id,
        lifecycleEventHookExecutionId=hook_execution_id,
        status=status
    )
    return {'status': status}

The Lambda hook function must have an IAM role with codedeploy:PutLifecycleEventHookExecutionStatus. If the hook times out (default 3600 seconds, configurable in the appspec) or returns Failed, the deployment fails and CodeDeploy rolls back to the blue target group. This gives you a programmatic gate: if the new MCP server version doesn't pass a health check against the test listener, it never receives production traffic.

Failure modes reference

FailureSymptomFix
ECS service update rejected"Service was unable to update the task definition" — ECS blocks the updateECS service is configured for ECS deployment controller, not CODE_DEPLOY; you cannot mix — service must be created with CODE_DEPLOY controller; recreate the service
Green tasks not passing health checksDeployment stuck at "Installation" phase; no traffic shift beginsALB health check on the green target group is failing; check the /health endpoint implementation on the new task definition; verify security group allows ALB to reach the container port
BeforeAllowTraffic hook times outDeployment stuck at "BeforeAllowTraffic" for hoursLambda hook did not call PutLifecycleEventHookExecutionStatus — check Lambda execution logs; common cause: Lambda VPC configuration prevents reaching the test listener, or IAM role missing codedeploy:PutLifecycleEventHookExecutionStatus
Rollback window closes before issue discoveredBlue tasks terminated; rollback requires full new deploymentIncrease terminationWaitTimeInMinutes (up to 2880 = 48 hours) and add CloudWatch alarms that trigger DEPLOYMENT_STOP_ON_ALARM; monitor error rates for at least one business day after each prod deploy
taskdef.json placeholder not substitutedappspec.yaml still contains literal string <TASK_DEFINITION>CodeDeploy only substitutes <TASK_DEFINITION> if taskdef.json is in the same artifact; verify your CodeBuild artifact includes both files; check for JSON syntax errors in taskdef.json that prevent registration
SSE MCP connections dropped during shiftMCP clients receive connection reset errors at traffic shift boundarySSE connections are long-lived; ALB connection draining setting (deregistration delay) on the blue target group controls how long the ALB keeps existing connections alive after deregistration — default is 300s; increase to 600s for long-running MCP tool calls