Guide · AWS CloudWatch · Observability

CloudWatch Logs Insights for MCP Server Debugging

CloudWatch Logs Insights gives you a SQL-like query engine over every Lambda log group — letting you trace individual MCP tool calls, pinpoint which tool names produce the most errors, and identify cold-start latency spikes across your entire MCP server fleet from a single interface. Unlike Grep-style log tailing, Logs Insights indexes each log event as a set of key-value fields, so structured JSON logs from AWS Lambda Powertools (or any JSON.stringify-based logger) become directly queryable without any preprocessing. For alarm-based alerting on the metrics you surface here, see CloudWatch Alarms for MCP servers; for container-level metrics when hosting on ECS/Fargate, see CloudWatch Container Insights for MCP servers.

TL;DR

Emit structured JSON logs from your MCP tool handlers using console.log(JSON.stringify({level:'INFO', tool:'s3_read', duration_ms: 142, ...})) or Lambda Powertools Logger. In CloudWatch Logs Insights, select the log groups for your MCP functions and use fields @timestamp, tool, duration_ms, errorMessage | filter level='ERROR' | sort @timestamp desc | limit 50 to see all tool failures in the last hour. Use stats avg(duration_ms), p99(duration_ms) by tool to rank your slowest MCP tools and prioritize cold-start fixes.

Why Logs Insights is the right debugging primitive for MCP servers

A production MCP server on Lambda generates three categories of logs: the Lambda platform report lines (REPORT, START, END), your structured application logs (tool call outcomes, durations, upstream API errors), and the MCP SDK's own debug output. All three land in the same log group, and standard CloudWatch Logs search can only do text match within individual log streams. Logs Insights ingests all streams from a log group simultaneously and builds an index that lets you aggregate, filter, and compute statistics across millions of log events in seconds.

The practical difference for MCP debugging: instead of guessing which Lambda execution environment produced a particular error and tailing four separate log streams, you write one query against /aws/lambda/mcp-server-prod and get all error events across all concurrent invocations sorted by time. You can then extend the query to group errors by MCP tool name, by the client user-agent that initiated the request, or by the upstream AWS service that returned an error — without leaving the CloudWatch console or writing a single line of log-processing code.

Debugging taskWithout Logs InsightsWith Logs Insights
Find all tool call errors in the last hourOpen each log stream individually, text-search for ERROROne query: filter level='ERROR' | sort @timestamp desc
Which tool name has the highest error rate?Manually tally errors from multiple streamsstats count(*) as errors by tool | sort errors desc
p99 latency for a specific toolExtract duration values, sort in a spreadsheetfilter tool='s3_list' | stats p99(duration_ms)
Identify cold starts vs warm invocationsSearch for "Init Duration" text across streamsfilter @message like 'Init Duration' | stats count(*) as cold_starts by bin(5m)
Trace a single MCP request end-to-endSearch for request ID across multiple log groupsMulti-group query: fields @logStream, @message | filter @requestId='abc-123'

Structuring MCP server logs for Logs Insights

Logs Insights automatically parses JSON log lines — any line that begins with { and ends with } is treated as a JSON object whose keys become queryable fields. You do not need to configure anything; the parsing happens at query time. This means you can start querying structured fields from existing JSON logs immediately, even logs that were emitted before you knew about Logs Insights.

The fields you choose to include in each log line determine what questions you can answer later. For an MCP server tool handler, a useful minimum log structure is:

// Minimal structured log for an MCP tool handler (TypeScript)
// Every log line is one JSON object — Logs Insights indexes each key as a field

interface McpToolLog {
  level: 'DEBUG' | 'INFO' | 'WARN' | 'ERROR';
  tool: string;           // The MCP tool name, e.g. "s3_read", "dynamo_query"
  duration_ms: number;    // Wall-clock time for the tool handler
  cold_start: boolean;    // true if this invocation initialized the Lambda container
  aws_request_id: string; // From context.awsRequestId — ties to the Lambda REPORT line
  mcp_session_id?: string; // Optional: the MCP session ID from the transport headers
  error?: string;          // Error message if the tool failed
  error_code?: string;     // Short error class: "ThrottlingException", "ValidationError", etc.
  upstream_service?: string; // e.g. "S3", "DynamoDB", "SecretsManager"
  input_size_bytes?: number; // Size of the tool input JSON
  output_size_bytes?: number; // Size of the tool output JSON
}

// Usage in a tool handler
async function s3ReadHandler(input: unknown, context: LambdaContext): Promise {
  const start = Date.now();
  try {
    const result = await s3Client.send(new GetObjectCommand({ Bucket: '...', Key: '...' }));
    const duration_ms = Date.now() - start;
    console.log(JSON.stringify({
      level: 'INFO',
      tool: 's3_read',
      duration_ms,
      cold_start: isFirstInvocation,
      aws_request_id: context.awsRequestId,
      output_size_bytes: result.ContentLength ?? 0,
    }));
    return { content: [{ type: 'text', text: await result.Body?.transformToString() ?? '' }] };
  } catch (err: unknown) {
    const duration_ms = Date.now() - start;
    const error = err as Error & { name?: string };
    console.log(JSON.stringify({
      level: 'ERROR',
      tool: 's3_read',
      duration_ms,
      cold_start: isFirstInvocation,
      aws_request_id: context.awsRequestId,
      error: error.message,
      error_code: error.name ?? 'UnknownError',
      upstream_service: 'S3',
    }));
    throw err;
  }
}

If you are using AWS Lambda Powertools for TypeScript, the Logger utility produces this structure automatically. Every logger.info() call adds cold_start, function_arn, function_name, function_request_id, and service to every log line. You only need to add your domain-specific fields like tool and duration_ms:

// Lambda Powertools Logger — structured JSON logging with built-in cold_start field
import { Logger } from '@aws-lambda-powertools/logger';

const logger = new Logger({ serviceName: 'mcp-server', logLevel: 'INFO' });

// In your handler:
logger.info('tool invocation', {
  tool: 's3_read',
  duration_ms: 142,
  output_size_bytes: 4096,
});

The emitted JSON will look like:

{
  "level": "INFO",
  "message": "tool invocation",
  "service": "mcp-server",
  "cold_start": false,
  "function_name": "mcp-server-prod",
  "function_request_id": "1a2b3c4d-...",
  "xray_trace_id": "1-66f3a2b1-...",
  "tool": "s3_read",
  "duration_ms": 142,
  "output_size_bytes": 4096,
  "timestamp": "2026-10-10T09:15:30.123Z"
}

Every key in that object becomes a queryable field in Logs Insights with no additional setup.

Essential Logs Insights queries for MCP servers

The following queries cover the most common MCP server debugging and monitoring scenarios. All queries assume your log group is /aws/lambda/mcp-server-prod and your logs contain a tool field. Adjust field names to match your actual log structure.

Error rate by tool name — last hour

fields @timestamp, tool, level, error, aws_request_id
| filter level = "ERROR"
| stats count(*) as error_count by tool
| sort error_count desc
| limit 20

p50, p90, p99 latency per tool — last 24 hours

fields @timestamp, tool, duration_ms
| filter ispresent(duration_ms) and ispresent(tool)
| stats
    percentile(duration_ms, 50) as p50_ms,
    percentile(duration_ms, 90) as p90_ms,
    percentile(duration_ms, 99) as p99_ms,
    count(*) as invocations
  by tool
| sort p99_ms desc

Cold start frequency — 5-minute buckets

fields @timestamp, cold_start
| filter cold_start = true
| stats count(*) as cold_starts by bin(5m)
| sort @timestamp asc

Lambda platform REPORT — extract Init Duration for cold starts

fields @timestamp, @message
| filter @message like /Init Duration/
| parse @message "Init Duration: * ms" as init_duration_ms
| stats
    avg(init_duration_ms) as avg_init_ms,
    max(init_duration_ms) as max_init_ms,
    count(*) as cold_start_count
  by bin(1h)

All errors from a specific upstream service

fields @timestamp, tool, error, error_code, aws_request_id
| filter upstream_service = "DynamoDB" and level = "ERROR"
| sort @timestamp desc
| limit 50

Throughput trend — invocations per minute

fields @timestamp, tool
| filter ispresent(tool)
| stats count(*) as invocations by bin(1m), tool
| sort @timestamp asc

Top callers by MCP session ID — identify heavy users

fields @timestamp, mcp_session_id, tool
| filter ispresent(mcp_session_id)
| stats count(*) as tool_calls by mcp_session_id
| sort tool_calls desc
| limit 20

Large output detection — find oversized tool responses

fields @timestamp, tool, output_size_bytes, aws_request_id
| filter output_size_bytes > 1000000
| sort output_size_bytes desc
| limit 20
Query patternWhen to use itKey field required
Error rate by toolIncident triage — find which tool is failinglevel, tool
Latency percentiles by toolPerformance baseline, SLO trackingduration_ms, tool
Cold start frequencyDiagnose tail latency spikes on low-traffic MCP serverscold_start
REPORT Init DurationMeasure actual initialization cost for SnapStart planningLambda platform log
Upstream service errorsCorrelate MCP tool failures with AWS service issuesupstream_service
Large output detectionFind tools that might hit MCP context window limitsoutput_size_bytes

Multi-log-group queries for distributed MCP servers

If your MCP server uses the pattern of one Lambda function per tool (common for CDK-built servers), you will have multiple log groups: /aws/lambda/mcp-tool-s3, /aws/lambda/mcp-tool-dynamo, /aws/lambda/mcp-tool-secrets, and so on. Logs Insights can query up to 50 log groups simultaneously using the group selector at the top of the console, or via the API with the logGroupNames array parameter.

To compare latency across all tool functions in one query:

# CLI: query multiple log groups simultaneously
aws logs start-query \
  --log-group-names \
    "/aws/lambda/mcp-tool-s3" \
    "/aws/lambda/mcp-tool-dynamo" \
    "/aws/lambda/mcp-tool-secrets" \
  --start-time $(date -d '1 hour ago' +%s) \
  --end-time $(date +%s) \
  --query-string '
    fields @logGroup, tool, duration_ms
    | filter ispresent(duration_ms)
    | stats
        percentile(duration_ms, 99) as p99_ms,
        count(*) as invocations
      by @logGroup
    | sort p99_ms desc
  '

# Poll for results
aws logs get-query-results --query-id $(jq -r .queryId query-result.json)

Saving queries and creating dashboards

Logs Insights queries can be saved as named favorites in the console (Actions → Save query) and referenced in CloudWatch Dashboards as Logs Insights widgets. This is useful for an MCP server ops dashboard: put a p99 latency time-series widget alongside a "top errors by tool" table widget and a cold-start frequency bar chart in a single dashboard that your team can bookmark.

# Create a CloudWatch Dashboard with a Logs Insights widget via AWS CLI
aws cloudwatch put-dashboard \
  --dashboard-name MCP-Server-Ops \
  --dashboard-body '{
    "widgets": [
      {
        "type": "log",
        "x": 0,
        "y": 0,
        "width": 24,
        "height": 6,
        "properties": {
          "title": "MCP Tool Error Rate",
          "region": "us-east-1",
          "logGroupNames": ["/aws/lambda/mcp-server-prod"],
          "query": "fields @timestamp, tool, level, error\n| filter level = \"ERROR\"\n| stats count(*) as errors by tool\n| sort errors desc",
          "view": "table"
        }
      }
    ]
  }'

AliveMCP uptime monitoring + Logs Insights: complementary layers

CloudWatch Logs Insights answers "what happened and why" — it is a post-hoc debugging and analysis tool. AliveMCP answers "is the server alive right now" — it probes your MCP endpoint from outside every 60 seconds using a real JSON-RPC tools/list handshake, not just an HTTP 200 check. The two are complementary: AliveMCP catches external reachability failures (DNS, TLS, cold-start timeouts visible to callers) that never produce a Lambda log line, while Logs Insights surfaces the internal error patterns that AliveMCP's probes don't reveal. Wire both: AliveMCP for real-user-visible uptime; Logs Insights + CloudWatch Alarms for internal health trending.

Failure modes

SymptomCauseFix
Query returns no results even though logs existTime range does not overlap with when logs were ingested, or log group name is misspelledStart with "Last 1 hour" and verify the log group name with aws logs describe-log-groups --log-group-name-prefix /aws/lambda/mcp
Fields like tool show as - (undefined)The log line is not valid JSON (e.g., a plain text error message from the MCP SDK or Node.js uncaught exception), so Logs Insights cannot parse the fieldSearch for those lines with filter NOT ispresent(tool) to see what plain-text log lines look like; wrap uncaught exceptions in a try/catch that logs structured JSON before re-throwing
Query is slow or times out on large time rangesLogs Insights scans all compressed log events in the requested time range — very large log volumes (>10GB) can take 30+ secondsNarrow the time range first; use filter level = 'ERROR' early in the query to reduce scanned data; consider using CloudWatch Metric Filters for high-cardinality metrics that you query frequently
Saved query no longer works after restructuring logsA field name changed (e.g., duration_ms renamed to durationMs) in a code refactorUse consistent snake_case field names and enforce them with a shared logging utility; version your log schema if you rename fields
Cross-account log query not workingCloudWatch Logs Insights does not natively cross AWS account boundariesUse CloudWatch Logs cross-account observability (requires setting up a monitoring account and source account sink policy) or use a log aggregation strategy — ship logs to a central log group via Kinesis Firehose and query there