Guide · AWS AppSync · Lambda Resolvers
AppSync Lambda Resolvers for MCP Servers — DIRECT vs BATCH Invoke, Error Shape, and the 30-Second Timeout
AppSync Lambda resolvers use an IAM service role to invoke a Lambda function as a data source, passing the GraphQL field arguments, identity, and source object in a structured event. There are two invocation modes: DIRECT (default — AppSync sends a single event per resolver invocation) and BATCH (AppSync accumulates multiple resolver invocations into a single Lambda call with an array of events — useful when a list field's child resolvers would otherwise fan out to N separate Lambda invocations). For MCP servers, Lambda resolvers are the right choice when the resolver logic is too complex for VTL or JavaScript resolver mapping — multi-step validation, calls to external APIs, or tool dispatch requiring dynamic routing. The hardest constraint: AppSync enforces a 30-second resolver timeout regardless of the Lambda function's own configured timeout. If the Lambda takes longer than 30 seconds, AppSync cancels the resolver and returns an error to the caller, but the Lambda continues running in the background — so tool operations that take up to 30 seconds must be idempotent to avoid double-execution on client retry.
TL;DR
Use DIRECT invoke for unit resolvers where one Lambda call = one GraphQL field. Use BATCH invoke when a list query resolves child fields via Lambda — AppSync groups invocations and sends a single { batchInvoke: true, source: [{…}] } event to Lambda, and Lambda must return an array of results in the same order. The Lambda's IAM execution role must trust appsync.amazonaws.com as a principal. The AppSync-imposed 30-second ceiling is independent of Lambda timeout — design long-running tool invocations as async (mutation → immediate response → poll/subscribe for completion). Response shape must exactly match the GraphQL field type; extra fields are silently ignored, but missing non-nullable fields cause a resolver error.
DIRECT invoke: event shape and response contract
In DIRECT mode, AppSync calls the Lambda once per resolver invocation. The event contains the full GraphQL context: arguments, identity, source (the parent object for nested resolvers), request headers, and info (field name, parent type, selection set). The Lambda's return value becomes the resolver result — it is mapped directly to the GraphQL field type without any VTL or JavaScript response mapping.
// Lambda handler — DIRECT invoke mode
// AppSync sends this event structure for a DIRECT Lambda data source
export const handler = async (event) => {
// event shape:
// {
// arguments: { serverId: "srv-123", includeHistory: true },
// identity: {
// sub: "user-uuid", // Cognito
// claims: { … }, // full JWT claims
// cognitoGroups: ["authors"]
// },
// source: null, // null for top-level query fields
// // non-null for nested resolvers (parent object)
// request: {
// headers: { "x-forwarded-for": "1.2.3.4", … }
// },
// info: {
// fieldName: "getMcpServerStatus",
// parentTypeName: "Query",
// variables: {},
// selectionSetList: ["serverId", "status", "lastCheckedAt", "uptimePct"]
// },
// prev: null // output of previous pipeline function (pipeline resolvers only)
// }
const { serverId, includeHistory } = event.arguments;
// Fetch server status
const status = await db.getServerStatus(serverId);
if (!status) {
// Returning null for a nullable field is valid
// For non-nullable fields, throw to propagate a GraphQL error
return null;
}
return {
serverId: status.serverId,
status: status.isHealthy ? "healthy" : "down",
lastCheckedAt: status.lastCheckedAt, // must be ISO 8601 for AWSDateTime fields
uptimePct: status.uptimePct,
history: includeHistory ? status.history : undefined // undefined = omit from response
};
// Extra fields are ignored; missing non-nullable fields cause a resolver error
};
// For GraphQL errors (caller sees errorMessage in the errors array):
// throw new Error("message") — errorType will be "Lambda:Unhandled"
// throw { message: "msg", errorType: "ServerOffline", data: {…} }
// The thrown object's message, errorType, and data are forwarded to the GraphQL error
BATCH invoke: accumulating N resolver calls into one Lambda invocation
When a GraphQL query returns a list of objects, and each object has a field resolved by the same Lambda data source, AppSync makes N separate Lambda invocations by default (one per list item). BATCH invoke changes this: AppSync waits for up to maxBatchSize items (default 200, max 2000), bundles them into a single Lambda event with { batchInvoke: true, source: [ item1, item2, … ] }, and expects the Lambda to return an array of results in the same order. This is the AppSync equivalent of a DataLoader pattern.
// Lambda handler — BATCH invoke mode
// AppSync accumulates up to maxBatchSize items and sends one event
export const handler = async (event) => {
// BATCH event shape:
// [
// { arguments: { field: "value" }, source: { serverId: "srv-1" }, identity: {…}, … },
// { arguments: { field: "value" }, source: { serverId: "srv-2" }, identity: {…}, … },
// …
// ]
// AppSync wraps the array in a list — handler receives the array directly
// Detect DIRECT vs BATCH from event structure
if (!Array.isArray(event)) {
// DIRECT invoke — handle as single item
return await resolveOne(event);
}
// BATCH invoke — event is an array of contexts
const serverIds = event.map(ctx => ctx.source.serverId);
// Batch fetch from DynamoDB using BatchGetItem
const results = await db.batchGetServerStatuses(serverIds);
// CRITICAL: return array in SAME ORDER as input
// If an item failed, return an error object at that position, not null
return event.map((ctx, i) => {
const status = results[ctx.source.serverId];
if (!status) {
// Return an error at this position — AppSync maps it to GraphQL error for this item
return {
data: null,
errorMessage: `Server ${ctx.source.serverId} not found`,
errorType: "NotFoundError"
};
}
return {
data: {
status: status.isHealthy ? "healthy" : "down",
lastCheckedAt: status.lastCheckedAt,
uptimePct: status.uptimePct
}
};
});
// Each array item must be either:
// - The plain result object (for success)
// - { data: result, errorMessage: "...", errorType: "..." } for partial failure
// Returning null at a position results in a null field (valid for nullable fields)
};
// CDK: enable BATCH on the LambdaDataSource
const ds = new appsync.LambdaDataSource(this, 'StatusDs', {
api,
lambdaFunction: statusLambda
});
new appsync.Resolver(this, 'StatusResolver', {
api,
typeName: 'Server',
fieldName: 'status',
dataSource: ds,
maxBatchSize: 100 // enables BATCH mode; 0 = DIRECT (default)
});
The order invariant in BATCH mode is absolute: AppSync maps the Lambda's response array back to the original request positions by index. If you return 3 items for a 4-item batch (for example, by filtering out a failed fetch), AppSync will silently assign wrong results to wrong objects. Always return an array of the exact same length as the input, using the { data, errorMessage, errorType } shape for failures.
IAM execution role and service principal
The AppSync Lambda data source requires an IAM service role that AppSync can assume to invoke the Lambda function. The role's trust policy must list appsync.amazonaws.com as the principal. This is separate from the Lambda function's resource-based policy — while you can also add a Lambda resource-based policy granting appsync.amazonaws.com, the service role approach is required for AppSync to call the Lambda across accounts or when the Lambda has no resource-based policy.
// CDK: create service role for AppSync → Lambda invocation
import * as iam from 'aws-cdk-lib/aws-iam';
const appSyncLambdaRole = new iam.Role(this, 'AppSyncLambdaRole', {
assumedBy: new iam.ServicePrincipal('appsync.amazonaws.com'),
inlinePolicies: {
InvokeLambda: new iam.PolicyDocument({
statements: [
new iam.PolicyStatement({
actions: ['lambda:InvokeFunction'],
resources: [
toolDispatchLambda.functionArn,
// Include :* for alias/version invocations if needed
]
})
]
})
}
});
const lambdaDs = new appsync.LambdaDataSource(this, 'ToolLambdaDs', {
api,
lambdaFunction: toolDispatchLambda,
serviceRole: appSyncLambdaRole // explicit role; CDK creates one automatically if omitted
});
// If using CDK auto-created role, verify it has lambda:InvokeFunction:
// lambdaDs.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({
// actions: ['lambda:InvokeFunction'],
// resources: [toolDispatchLambda.functionArn]
// }));
The 30-second resolver timeout: async tool pattern
AppSync enforces a hard 30-second ceiling on resolver execution, independent of the Lambda's own configured timeout. A tool that takes 45 seconds results in: AppSync returns a resolver timeout error to the caller at 30 seconds, but the Lambda continues running in the background for another 15 seconds. The safe pattern for long-running MCP tools is to return immediately with a job ID, then let the client subscribe for the result.
// Async tool dispatch pattern — return job ID immediately
// 1. Mutation: start job, return jobId
// 2. AppSync subscription: subscribe for onToolComplete(jobId: "...")
// 3. Job Lambda: does the work, writes result, triggers createToolResult mutation
// Mutation resolver (completes in < 1 second)
export const startToolHandler = async (event) => {
const { toolName, sessionId, toolInput } = event.arguments;
// Write a pending job record
const jobId = crypto.randomUUID();
await db.createToolJob({ jobId, toolName, sessionId, toolInput, status: "pending" });
// Enqueue to SQS for async processing (not a direct Lambda call)
await sqs.send(new SendMessageCommand({
QueueUrl: process.env.TOOL_QUEUE_URL,
MessageBody: JSON.stringify({ jobId, toolName, toolInput, sessionId })
}));
// Return immediately — well within 30 seconds
return {
jobId,
status: "pending",
createdAt: new Date().toISOString()
};
};
// Background tool worker Lambda (triggered by SQS, NOT by AppSync)
// This Lambda can run for up to 15 minutes (or configured timeout)
export const toolWorkerHandler = async (event) => {
const { jobId, toolName, toolInput, sessionId } = JSON.parse(event.Records[0].body);
const result = await runTool(toolName, toolInput); // may take minutes
// Write result to AppSync via GraphQL mutation (IAM-signed request)
// This triggers the subscription: onToolComplete(jobId: jobId)
await appsyncMutation(`
mutation CompleteToolJob($input: ToolJobCompleteInput!) {
completeToolJob(input: $input) { jobId status output durationMs }
}
`, { input: { jobId, status: "complete", output: JSON.stringify(result), durationMs: result.durationMs } });
};
Failure modes reference
| Failure | Symptom | Fix |
|---|---|---|
| Lambda throws unhandled exception | GraphQL error with errorType: "Lambda:Unhandled" and the exception message | Catch all errors and return structured error objects; throw { message, errorType, data } for typed GraphQL errors |
| BATCH response array shorter than input array | Wrong results silently mapped to wrong parent objects | Always return array of exact same length as input; use { data: null, errorMessage, errorType } for missing items |
| AppSync 30-second timeout — Lambda still running | Caller gets resolver timeout error; Lambda continues in background; duplicate side effects on retry | Design mutations to be idempotent (deduplicate by clientMutationId); use async pattern (job queue + subscription) for operations > 5 seconds |
Service role missing lambda:InvokeFunction | All resolver invocations fail with Access denied at the data source call stage | Grant lambda:InvokeFunction on the Lambda ARN to the AppSync service role (appsync.amazonaws.com principal) |
| Non-nullable field missing in Lambda return value | Resolver error: "Cannot return null for non-nullable type" | Include all non-nullable fields in the return object; use undefined (not null) for absent optional fields — AppSync omits undefined values from the response |
| DIRECT event received in BATCH handler (or vice versa) | TypeError: event.map is not a function | Check Array.isArray(event) to distinguish BATCH vs DIRECT events; handle both in the same Lambda or use separate Lambdas per mode |
| Lambda cold start exceeds 30 seconds under BATCH with large items | First batch after idle period times out; subsequent batches succeed | Enable Lambda provisioned concurrency for AppSync data source Lambdas with latency-sensitive resolvers |