Guide · AWS CloudWatch · ECS / Fargate

CloudWatch Container Insights for ECS/Fargate MCP Servers

CloudWatch Container Insights collects CPU utilization, memory usage, network I/O, and storage metrics at the ECS task and container level — without any changes to your MCP server application code — and publishes them to the ECS/ContainerInsights namespace where you can alarm, dashboard, and query them alongside your Lambda metrics. For Fargate-hosted MCP servers, this is the primary resource-utilization observability layer: it tells you when a task is hitting its memory limit before the OOM kill, when CPU is throttled under sustained tool-call load, and whether your task definitions need a resource adjustment. For alarm configuration on top of these metrics, see CloudWatch Alarms for MCP servers; for log-level query analysis, see CloudWatch Logs Insights for MCP servers.

TL;DR

Enable Container Insights on your ECS cluster with aws ecs update-cluster-settings --cluster mcp-cluster --settings name=containerInsights,value=enabled. Metrics appear in the ECS/ContainerInsights CloudWatch namespace within a few minutes. Key metrics to watch: CpuUtilized and MemoryUtilized (absolute values against CpuReserved / MemoryReserved) and TaskCount to confirm service scaling is working. On Fargate, Container Insights collects metrics automatically with no sidecar required.

Why Container Insights matters for MCP server deployments

When an MCP server runs on Lambda, AWS publishes Duration, Errors, and Throttles metrics automatically. When the same server runs on ECS/Fargate as a long-lived Node.js or Python process, those Lambda metrics do not exist — you only get container-level metrics. Without Container Insights enabled, you have no CloudWatch data about how much CPU the MCP server container is using, whether it is approaching its memory limit, or how much network traffic the tool calls are generating.

Container Insights fills this gap. It runs a CloudWatch agent inside the ECS agent (on EC2 launch type) or as part of the Fargate platform (on Fargate launch type) that samples resource utilization from the container runtime and publishes it to CloudWatch with dimensions for cluster, service, task definition, and individual container. The result is a per-container, per-task metric set that you can alarm on, visualize, and query — without touching your MCP server application code.

MetricNamespaceWhat it tells you about your MCP server
CpuUtilizedECS/ContainerInsightsvCPU units consumed by the task; compare to CpuReserved to see CPU utilization %
MemoryUtilizedECS/ContainerInsightsMiB of memory in use; approaching MemoryReserved means risk of OOM kill
NetworkRxBytes / NetworkTxBytesECS/ContainerInsightsBytes received/sent by the task; spikes correlate with large tool call payloads
StorageReadBytes / StorageWriteBytesECS/ContainerInsightsEphemeral storage I/O; relevant if tools write to the container filesystem
TaskCountECS/ContainerInsightsRunning task count; confirms service autoscaling is working as expected
RunningTaskCountECS/ContainerInsightsAt service granularity; drop to 0 = service is down, MCP endpoint unreachable

Enabling Container Insights on an existing ECS cluster

# Enable Container Insights on an existing cluster
# This change is non-destructive and takes effect immediately for new tasks
aws ecs update-cluster-settings \
  --cluster mcp-cluster \
  --settings name=containerInsights,value=enabled

# Verify the setting was applied
aws ecs describe-clusters \
  --clusters mcp-cluster \
  --query 'clusters[0].settings'
# Expected output:
# [{"name": "containerInsights", "value": "enabled"}]

# For new cluster creation (recommended for new deployments)
aws ecs create-cluster \
  --cluster-name mcp-cluster \
  --settings name=containerInsights,value=enabled

For Fargate launch type, no additional configuration is required after enabling Container Insights on the cluster. Metrics will appear in CloudWatch within 5–10 minutes of the next task start. For EC2 launch type, the ECS-optimized AMI ships the CloudWatch agent pre-installed; ensure your EC2 instances have the CloudWatchAgentServerPolicy IAM policy attached to their instance profile.

CDK: enable Container Insights on cluster creation

import * as ecs from 'aws-cdk-lib/aws-ecs';
import * as cdk from 'aws-cdk-lib';
import { Construct } from 'constructs';

export class McpCluster extends Construct {
  public readonly cluster: ecs.Cluster;

  constructor(scope: Construct, id: string) {
    super(scope, id);

    this.cluster = new ecs.Cluster(this, 'McpCluster', {
      clusterName: 'mcp-cluster',
      containerInsights: true,  // enables ECS/ContainerInsights metrics
      vpc: /* your VPC */,
    });
  }
}

Querying Container Insights metrics

Container Insights metrics live in the ECS/ContainerInsights namespace. The primary dimension combinations available are: ClusterName only (cluster-aggregate), ClusterName + ServiceName (service-level), and ClusterName + TaskDefinitionFamily (task definition level). Container-level granularity is available in the pre-built Container Insights console dashboards but not directly as CloudWatch metrics — it is written to a special performance log stream in CloudWatch Logs which you can query with Logs Insights.

# Get CPU utilization for an MCP server service over the last hour
aws cloudwatch get-metric-statistics \
  --namespace "ECS/ContainerInsights" \
  --metric-name "CpuUtilized" \
  --dimensions \
    Name=ClusterName,Value=mcp-cluster \
    Name=ServiceName,Value=mcp-server-service \
  --start-time $(date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%SZ) \
  --end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
  --period 300 \
  --statistics Average Maximum \
  --query 'sort_by(Datapoints, &Timestamp)[*].{Time:Timestamp,Avg:Average,Max:Maximum}'

# Calculate CPU utilization percentage
# CpuUtilized is in vCPU units; CpuReserved is the task definition CPU value
# Utilization % = (CpuUtilized / CpuReserved) * 100
# Both metrics share the same dimensions — you need to fetch both and divide

Container-level metrics via Logs Insights

For per-container breakdown (useful when a task definition has a sidecar — e.g., an MCP server container + a log router container), the container-granularity data is in the /aws/ecs/containerinsights/mcp-cluster/performance log group as structured JSON. Query it with Logs Insights:

-- Query the Container Insights performance log group for container-level metrics
-- Log group: /aws/ecs/containerinsights/mcp-cluster/performance

fields @timestamp, ContainerName, CpuUtilized, MemoryUtilized, MemoryReserved
| filter Type = "Container" and ServiceName = "mcp-server-service"
| stats
    avg(CpuUtilized) as avg_cpu,
    max(MemoryUtilized) as max_mem,
    avg(MemoryReserved) as reserved_mem
  by ContainerName, bin(5m)
| sort @timestamp asc

Alarms on Container Insights metrics for MCP servers

The most operationally useful Container Insights alarms for an MCP server are on memory utilization (OOM risk) and running task count (service health). CPU alarms are secondary — ECS can throttle CPU but will not kill tasks for CPU overuse the way it does for memory overuse, so CPU alarms are indicators of performance degradation rather than availability risk.

# Alarm: MCP server memory utilization approaching limit
# ECS/ContainerInsights MemoryUtilized is in MiB; MemoryReserved is the task definition value
# Alarm when average memory is above 80% of reserved for 2 consecutive 5-minute windows

aws cloudwatch put-metric-alarm \
  --alarm-name "mcp-server-memory-high" \
  --alarm-description "MCP server task memory above 80% of reservation" \
  --metrics '[
    {
      "Id": "mem_used",
      "MetricStat": {
        "Metric": {
          "Namespace": "ECS/ContainerInsights",
          "MetricName": "MemoryUtilized",
          "Dimensions": [
            {"Name": "ClusterName", "Value": "mcp-cluster"},
            {"Name": "ServiceName", "Value": "mcp-server-service"}
          ]
        },
        "Period": 300,
        "Stat": "Average"
      }
    },
    {
      "Id": "mem_reserved",
      "MetricStat": {
        "Metric": {
          "Namespace": "ECS/ContainerInsights",
          "MetricName": "MemoryReserved",
          "Dimensions": [
            {"Name": "ClusterName", "Value": "mcp-cluster"},
            {"Name": "ServiceName", "Value": "mcp-server-service"}
          ]
        },
        "Period": 300,
        "Stat": "Average"
      }
    },
    {
      "Id": "utilization_pct",
      "Expression": "(mem_used / mem_reserved) * 100"
    }
  ]' \
  --comparison-operator GreaterThanOrEqualToThreshold \
  --threshold 80 \
  --threshold-metric-id utilization_pct \
  --evaluation-periods 2 \
  --treat-missing-data notBreaching \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts

# Alarm: MCP server has zero running tasks (service is down)
aws cloudwatch put-metric-alarm \
  --alarm-name "mcp-server-no-running-tasks" \
  --alarm-description "MCP server service has no running tasks" \
  --namespace "ECS/ContainerInsights" \
  --metric-name "RunningTaskCount" \
  --dimensions \
    Name=ClusterName,Value=mcp-cluster \
    Name=ServiceName,Value=mcp-server-service \
  --statistic Average \
  --period 60 \
  --evaluation-periods 2 \
  --threshold 1 \
  --comparison-operator LessThanThreshold \
  --treat-missing-data breaching \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:mcp-alerts-critical

Task definition sizing for MCP servers

Container Insights data is the right input for sizing MCP server task definitions. A common mistake is to allocate the same CPU and memory as the production Lambda memory limit (e.g., 512MB) without accounting for the fact that a Fargate task runs a full container — including the Node.js runtime, the MCP SDK's in-memory tool registry, and any connection pools your tools maintain. Check the MemoryUtilized p99 under realistic load and add at least a 30% headroom margin to your MemoryReservation to avoid OOM kills under burst traffic.

MCP server profileRecommended CPURecommended MemorySizing rationale
Lightweight (1–3 tools, no local state)0.25 vCPU (256 CPU units)512 MiBNode.js runtime ~80MB; MCP SDK ~30MB; tool code ~10MB; 30% headroom
Standard (5–10 tools, SQLite or Redis connection)0.5 vCPU (512 CPU units)1024 MiBConnection pool overhead; SQLite in-memory pages; concurrent tool call buffers
Heavy (10+ tools, LLM embedding calls, large payloads)1 vCPU (1024 CPU units)2048 MiBEmbedding model buffers; large tool output buffering; concurrent HTTP connections
AI-intensive (local model inference, vector operations)4 vCPU (4096 CPU units)8192 MiBModel weights in memory; ONNX runtime; batch processing

AliveMCP external probing + Container Insights

Container Insights tells you the container is consuming resources — but a container can be consuming CPU and memory while not correctly serving MCP requests. AliveMCP probes the actual MCP protocol endpoint from outside the ECS cluster every 60 seconds, testing the full network path and JSON-RPC handshake. If your MCP server container is running but stuck in an infinite loop (high CPU, no useful output) or has a hung connection pool (normal resource usage, no responses), Container Insights alone will not surface the problem. AliveMCP will. Use both: Container Insights for resource health, AliveMCP for endpoint reachability.

Failure modes

SymptomCauseFix
No metrics in ECS/ContainerInsights namespaceContainer Insights not enabled on the cluster, or tasks started before Container Insights was enabled (metrics only appear for tasks started after enabling)Enable Container Insights and restart the service: aws ecs update-service --cluster mcp-cluster --service mcp-server-service --force-new-deployment
RunningTaskCount is 0 but tasks appear running in consoleContainer Insights metrics are published every 60 seconds; there can be up to 2–3 minutes of lag after a deployment before the metric stabilizesWait 3–5 minutes; if still 0, check the ECS Agent logs on the EC2 host or verify Fargate platform version is 1.4+ (earlier versions had Container Insights gaps)
MemoryUtilized metric is always equal to MemoryReservedFargate rounds up MemoryUtilized to MemoryReserved when the actual usage is near the limit; this can mean the container is OOM-killing tasks without an alarmCheck ECS service events in the console for OutOfMemoryError task stop codes; increase MemoryReserved and monitor whether MemoryUtilized drops below the reservation
Container Insights data costs more than expectedContainer Insights metrics are published to CloudWatch at 1-minute resolution, which incurs CloudWatch Metrics charges (~$0.30/metric/month); a large ECS cluster with many services and task definitions generates many metric combinationsReview the full list of Container Insights metrics emitted with aws cloudwatch list-metrics --namespace ECS/ContainerInsights; disable Container Insights on dev clusters where you do not need it
Alarm on CPU never fires despite clearly throttled containersCpuUtilized is absolute vCPU units; if your alarm threshold is set to a raw value (e.g., 256 = 25% of a 1-vCPU task) but you have multiple tasks, the per-service Average stays below the threshold even though individual tasks are throttledUse a metric math expression that computes utilization percentage: (CpuUtilized / CpuReserved) * 100 and alarm on the percentage (e.g., >80%)