Guide · AWS CloudWatch · Observability
CloudWatch Logs Insights for MCP Server Debugging
CloudWatch Logs Insights gives you a SQL-like query engine over every Lambda log group — letting you trace individual MCP tool calls, pinpoint which tool names produce the most errors, and identify cold-start latency spikes across your entire MCP server fleet from a single interface. Unlike Grep-style log tailing, Logs Insights indexes each log event as a set of key-value fields, so structured JSON logs from AWS Lambda Powertools (or any JSON.stringify-based logger) become directly queryable without any preprocessing. For alarm-based alerting on the metrics you surface here, see CloudWatch Alarms for MCP servers; for container-level metrics when hosting on ECS/Fargate, see CloudWatch Container Insights for MCP servers.
TL;DR
Emit structured JSON logs from your MCP tool handlers using console.log(JSON.stringify({level:'INFO', tool:'s3_read', duration_ms: 142, ...})) or Lambda Powertools Logger. In CloudWatch Logs Insights, select the log groups for your MCP functions and use fields @timestamp, tool, duration_ms, errorMessage | filter level='ERROR' | sort @timestamp desc | limit 50 to see all tool failures in the last hour. Use stats avg(duration_ms), p99(duration_ms) by tool to rank your slowest MCP tools and prioritize cold-start fixes.
Why Logs Insights is the right debugging primitive for MCP servers
A production MCP server on Lambda generates three categories of logs: the Lambda platform report lines (REPORT, START, END), your structured application logs (tool call outcomes, durations, upstream API errors), and the MCP SDK's own debug output. All three land in the same log group, and standard CloudWatch Logs search can only do text match within individual log streams. Logs Insights ingests all streams from a log group simultaneously and builds an index that lets you aggregate, filter, and compute statistics across millions of log events in seconds.
The practical difference for MCP debugging: instead of guessing which Lambda execution environment produced a particular error and tailing four separate log streams, you write one query against /aws/lambda/mcp-server-prod and get all error events across all concurrent invocations sorted by time. You can then extend the query to group errors by MCP tool name, by the client user-agent that initiated the request, or by the upstream AWS service that returned an error — without leaving the CloudWatch console or writing a single line of log-processing code.
| Debugging task | Without Logs Insights | With Logs Insights |
|---|---|---|
| Find all tool call errors in the last hour | Open each log stream individually, text-search for ERROR | One query: filter level='ERROR' | sort @timestamp desc |
| Which tool name has the highest error rate? | Manually tally errors from multiple streams | stats count(*) as errors by tool | sort errors desc |
| p99 latency for a specific tool | Extract duration values, sort in a spreadsheet | filter tool='s3_list' | stats p99(duration_ms) |
| Identify cold starts vs warm invocations | Search for "Init Duration" text across streams | filter @message like 'Init Duration' | stats count(*) as cold_starts by bin(5m) |
| Trace a single MCP request end-to-end | Search for request ID across multiple log groups | Multi-group query: fields @logStream, @message | filter @requestId='abc-123' |
Structuring MCP server logs for Logs Insights
Logs Insights automatically parses JSON log lines — any line that begins with { and ends with } is treated as a JSON object whose keys become queryable fields. You do not need to configure anything; the parsing happens at query time. This means you can start querying structured fields from existing JSON logs immediately, even logs that were emitted before you knew about Logs Insights.
The fields you choose to include in each log line determine what questions you can answer later. For an MCP server tool handler, a useful minimum log structure is:
// Minimal structured log for an MCP tool handler (TypeScript)
// Every log line is one JSON object — Logs Insights indexes each key as a field
interface McpToolLog {
level: 'DEBUG' | 'INFO' | 'WARN' | 'ERROR';
tool: string; // The MCP tool name, e.g. "s3_read", "dynamo_query"
duration_ms: number; // Wall-clock time for the tool handler
cold_start: boolean; // true if this invocation initialized the Lambda container
aws_request_id: string; // From context.awsRequestId — ties to the Lambda REPORT line
mcp_session_id?: string; // Optional: the MCP session ID from the transport headers
error?: string; // Error message if the tool failed
error_code?: string; // Short error class: "ThrottlingException", "ValidationError", etc.
upstream_service?: string; // e.g. "S3", "DynamoDB", "SecretsManager"
input_size_bytes?: number; // Size of the tool input JSON
output_size_bytes?: number; // Size of the tool output JSON
}
// Usage in a tool handler
async function s3ReadHandler(input: unknown, context: LambdaContext): Promise {
const start = Date.now();
try {
const result = await s3Client.send(new GetObjectCommand({ Bucket: '...', Key: '...' }));
const duration_ms = Date.now() - start;
console.log(JSON.stringify({
level: 'INFO',
tool: 's3_read',
duration_ms,
cold_start: isFirstInvocation,
aws_request_id: context.awsRequestId,
output_size_bytes: result.ContentLength ?? 0,
}));
return { content: [{ type: 'text', text: await result.Body?.transformToString() ?? '' }] };
} catch (err: unknown) {
const duration_ms = Date.now() - start;
const error = err as Error & { name?: string };
console.log(JSON.stringify({
level: 'ERROR',
tool: 's3_read',
duration_ms,
cold_start: isFirstInvocation,
aws_request_id: context.awsRequestId,
error: error.message,
error_code: error.name ?? 'UnknownError',
upstream_service: 'S3',
}));
throw err;
}
}
If you are using AWS Lambda Powertools for TypeScript, the Logger utility produces this structure automatically. Every logger.info() call adds cold_start, function_arn, function_name, function_request_id, and service to every log line. You only need to add your domain-specific fields like tool and duration_ms:
// Lambda Powertools Logger — structured JSON logging with built-in cold_start field
import { Logger } from '@aws-lambda-powertools/logger';
const logger = new Logger({ serviceName: 'mcp-server', logLevel: 'INFO' });
// In your handler:
logger.info('tool invocation', {
tool: 's3_read',
duration_ms: 142,
output_size_bytes: 4096,
});
The emitted JSON will look like:
{
"level": "INFO",
"message": "tool invocation",
"service": "mcp-server",
"cold_start": false,
"function_name": "mcp-server-prod",
"function_request_id": "1a2b3c4d-...",
"xray_trace_id": "1-66f3a2b1-...",
"tool": "s3_read",
"duration_ms": 142,
"output_size_bytes": 4096,
"timestamp": "2026-10-10T09:15:30.123Z"
}
Every key in that object becomes a queryable field in Logs Insights with no additional setup.
Essential Logs Insights queries for MCP servers
The following queries cover the most common MCP server debugging and monitoring scenarios. All queries assume your log group is /aws/lambda/mcp-server-prod and your logs contain a tool field. Adjust field names to match your actual log structure.
Error rate by tool name — last hour
fields @timestamp, tool, level, error, aws_request_id
| filter level = "ERROR"
| stats count(*) as error_count by tool
| sort error_count desc
| limit 20
p50, p90, p99 latency per tool — last 24 hours
fields @timestamp, tool, duration_ms
| filter ispresent(duration_ms) and ispresent(tool)
| stats
percentile(duration_ms, 50) as p50_ms,
percentile(duration_ms, 90) as p90_ms,
percentile(duration_ms, 99) as p99_ms,
count(*) as invocations
by tool
| sort p99_ms desc
Cold start frequency — 5-minute buckets
fields @timestamp, cold_start
| filter cold_start = true
| stats count(*) as cold_starts by bin(5m)
| sort @timestamp asc
Lambda platform REPORT — extract Init Duration for cold starts
fields @timestamp, @message
| filter @message like /Init Duration/
| parse @message "Init Duration: * ms" as init_duration_ms
| stats
avg(init_duration_ms) as avg_init_ms,
max(init_duration_ms) as max_init_ms,
count(*) as cold_start_count
by bin(1h)
All errors from a specific upstream service
fields @timestamp, tool, error, error_code, aws_request_id
| filter upstream_service = "DynamoDB" and level = "ERROR"
| sort @timestamp desc
| limit 50
Throughput trend — invocations per minute
fields @timestamp, tool
| filter ispresent(tool)
| stats count(*) as invocations by bin(1m), tool
| sort @timestamp asc
Top callers by MCP session ID — identify heavy users
fields @timestamp, mcp_session_id, tool
| filter ispresent(mcp_session_id)
| stats count(*) as tool_calls by mcp_session_id
| sort tool_calls desc
| limit 20
Large output detection — find oversized tool responses
fields @timestamp, tool, output_size_bytes, aws_request_id
| filter output_size_bytes > 1000000
| sort output_size_bytes desc
| limit 20
| Query pattern | When to use it | Key field required |
|---|---|---|
| Error rate by tool | Incident triage — find which tool is failing | level, tool |
| Latency percentiles by tool | Performance baseline, SLO tracking | duration_ms, tool |
| Cold start frequency | Diagnose tail latency spikes on low-traffic MCP servers | cold_start |
| REPORT Init Duration | Measure actual initialization cost for SnapStart planning | Lambda platform log |
| Upstream service errors | Correlate MCP tool failures with AWS service issues | upstream_service |
| Large output detection | Find tools that might hit MCP context window limits | output_size_bytes |
Multi-log-group queries for distributed MCP servers
If your MCP server uses the pattern of one Lambda function per tool (common for CDK-built servers), you will have multiple log groups: /aws/lambda/mcp-tool-s3, /aws/lambda/mcp-tool-dynamo, /aws/lambda/mcp-tool-secrets, and so on. Logs Insights can query up to 50 log groups simultaneously using the group selector at the top of the console, or via the API with the logGroupNames array parameter.
To compare latency across all tool functions in one query:
# CLI: query multiple log groups simultaneously
aws logs start-query \
--log-group-names \
"/aws/lambda/mcp-tool-s3" \
"/aws/lambda/mcp-tool-dynamo" \
"/aws/lambda/mcp-tool-secrets" \
--start-time $(date -d '1 hour ago' +%s) \
--end-time $(date +%s) \
--query-string '
fields @logGroup, tool, duration_ms
| filter ispresent(duration_ms)
| stats
percentile(duration_ms, 99) as p99_ms,
count(*) as invocations
by @logGroup
| sort p99_ms desc
'
# Poll for results
aws logs get-query-results --query-id $(jq -r .queryId query-result.json)
Saving queries and creating dashboards
Logs Insights queries can be saved as named favorites in the console (Actions → Save query) and referenced in CloudWatch Dashboards as Logs Insights widgets. This is useful for an MCP server ops dashboard: put a p99 latency time-series widget alongside a "top errors by tool" table widget and a cold-start frequency bar chart in a single dashboard that your team can bookmark.
# Create a CloudWatch Dashboard with a Logs Insights widget via AWS CLI
aws cloudwatch put-dashboard \
--dashboard-name MCP-Server-Ops \
--dashboard-body '{
"widgets": [
{
"type": "log",
"x": 0,
"y": 0,
"width": 24,
"height": 6,
"properties": {
"title": "MCP Tool Error Rate",
"region": "us-east-1",
"logGroupNames": ["/aws/lambda/mcp-server-prod"],
"query": "fields @timestamp, tool, level, error\n| filter level = \"ERROR\"\n| stats count(*) as errors by tool\n| sort errors desc",
"view": "table"
}
}
]
}'
AliveMCP uptime monitoring + Logs Insights: complementary layers
CloudWatch Logs Insights answers "what happened and why" — it is a post-hoc debugging and analysis tool. AliveMCP answers "is the server alive right now" — it probes your MCP endpoint from outside every 60 seconds using a real JSON-RPC tools/list handshake, not just an HTTP 200 check. The two are complementary: AliveMCP catches external reachability failures (DNS, TLS, cold-start timeouts visible to callers) that never produce a Lambda log line, while Logs Insights surfaces the internal error patterns that AliveMCP's probes don't reveal. Wire both: AliveMCP for real-user-visible uptime; Logs Insights + CloudWatch Alarms for internal health trending.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Query returns no results even though logs exist | Time range does not overlap with when logs were ingested, or log group name is misspelled | Start with "Last 1 hour" and verify the log group name with aws logs describe-log-groups --log-group-name-prefix /aws/lambda/mcp |
Fields like tool show as - (undefined) | The log line is not valid JSON (e.g., a plain text error message from the MCP SDK or Node.js uncaught exception), so Logs Insights cannot parse the field | Search for those lines with filter NOT ispresent(tool) to see what plain-text log lines look like; wrap uncaught exceptions in a try/catch that logs structured JSON before re-throwing |
| Query is slow or times out on large time ranges | Logs Insights scans all compressed log events in the requested time range — very large log volumes (>10GB) can take 30+ seconds | Narrow the time range first; use filter level = 'ERROR' early in the query to reduce scanned data; consider using CloudWatch Metric Filters for high-cardinality metrics that you query frequently |
| Saved query no longer works after restructuring logs | A field name changed (e.g., duration_ms renamed to durationMs) in a code refactor | Use consistent snake_case field names and enforce them with a shared logging utility; version your log schema if you rename fields |
| Cross-account log query not working | CloudWatch Logs Insights does not natively cross AWS account boundaries | Use CloudWatch Logs cross-account observability (requires setting up a monitoring account and source account sink policy) or use a log aggregation strategy — ship logs to a central log group via Kinesis Firehose and query there |