Guide · AWS AppSync · Lambda Resolvers

AppSync Lambda Resolvers for MCP Servers — DIRECT vs BATCH Invoke, Error Shape, and the 30-Second Timeout

AppSync Lambda resolvers use an IAM service role to invoke a Lambda function as a data source, passing the GraphQL field arguments, identity, and source object in a structured event. There are two invocation modes: DIRECT (default — AppSync sends a single event per resolver invocation) and BATCH (AppSync accumulates multiple resolver invocations into a single Lambda call with an array of events — useful when a list field's child resolvers would otherwise fan out to N separate Lambda invocations). For MCP servers, Lambda resolvers are the right choice when the resolver logic is too complex for VTL or JavaScript resolver mapping — multi-step validation, calls to external APIs, or tool dispatch requiring dynamic routing. The hardest constraint: AppSync enforces a 30-second resolver timeout regardless of the Lambda function's own configured timeout. If the Lambda takes longer than 30 seconds, AppSync cancels the resolver and returns an error to the caller, but the Lambda continues running in the background — so tool operations that take up to 30 seconds must be idempotent to avoid double-execution on client retry.

TL;DR

Use DIRECT invoke for unit resolvers where one Lambda call = one GraphQL field. Use BATCH invoke when a list query resolves child fields via Lambda — AppSync groups invocations and sends a single { batchInvoke: true, source: [{…}] } event to Lambda, and Lambda must return an array of results in the same order. The Lambda's IAM execution role must trust appsync.amazonaws.com as a principal. The AppSync-imposed 30-second ceiling is independent of Lambda timeout — design long-running tool invocations as async (mutation → immediate response → poll/subscribe for completion). Response shape must exactly match the GraphQL field type; extra fields are silently ignored, but missing non-nullable fields cause a resolver error.

DIRECT invoke: event shape and response contract

In DIRECT mode, AppSync calls the Lambda once per resolver invocation. The event contains the full GraphQL context: arguments, identity, source (the parent object for nested resolvers), request headers, and info (field name, parent type, selection set). The Lambda's return value becomes the resolver result — it is mapped directly to the GraphQL field type without any VTL or JavaScript response mapping.

// Lambda handler — DIRECT invoke mode
// AppSync sends this event structure for a DIRECT Lambda data source
export const handler = async (event) => {
  // event shape:
  // {
  //   arguments: { serverId: "srv-123", includeHistory: true },
  //   identity: {
  //     sub: "user-uuid",        // Cognito
  //     claims: { … },           // full JWT claims
  //     cognitoGroups: ["authors"]
  //   },
  //   source: null,              // null for top-level query fields
  //                              // non-null for nested resolvers (parent object)
  //   request: {
  //     headers: { "x-forwarded-for": "1.2.3.4", … }
  //   },
  //   info: {
  //     fieldName: "getMcpServerStatus",
  //     parentTypeName: "Query",
  //     variables: {},
  //     selectionSetList: ["serverId", "status", "lastCheckedAt", "uptimePct"]
  //   },
  //   prev: null                 // output of previous pipeline function (pipeline resolvers only)
  // }

  const { serverId, includeHistory } = event.arguments;

  // Fetch server status
  const status = await db.getServerStatus(serverId);
  if (!status) {
    // Returning null for a nullable field is valid
    // For non-nullable fields, throw to propagate a GraphQL error
    return null;
  }

  return {
    serverId: status.serverId,
    status: status.isHealthy ? "healthy" : "down",
    lastCheckedAt: status.lastCheckedAt,  // must be ISO 8601 for AWSDateTime fields
    uptimePct: status.uptimePct,
    history: includeHistory ? status.history : undefined  // undefined = omit from response
  };
  // Extra fields are ignored; missing non-nullable fields cause a resolver error
};

// For GraphQL errors (caller sees errorMessage in the errors array):
// throw new Error("message") — errorType will be "Lambda:Unhandled"
// throw { message: "msg", errorType: "ServerOffline", data: {…} }
// The thrown object's message, errorType, and data are forwarded to the GraphQL error

BATCH invoke: accumulating N resolver calls into one Lambda invocation

When a GraphQL query returns a list of objects, and each object has a field resolved by the same Lambda data source, AppSync makes N separate Lambda invocations by default (one per list item). BATCH invoke changes this: AppSync waits for up to maxBatchSize items (default 200, max 2000), bundles them into a single Lambda event with { batchInvoke: true, source: [ item1, item2, … ] }, and expects the Lambda to return an array of results in the same order. This is the AppSync equivalent of a DataLoader pattern.

// Lambda handler — BATCH invoke mode
// AppSync accumulates up to maxBatchSize items and sends one event
export const handler = async (event) => {
  // BATCH event shape:
  // [
  //   { arguments: { field: "value" }, source: { serverId: "srv-1" }, identity: {…}, … },
  //   { arguments: { field: "value" }, source: { serverId: "srv-2" }, identity: {…}, … },
  //   …
  // ]
  // AppSync wraps the array in a list — handler receives the array directly

  // Detect DIRECT vs BATCH from event structure
  if (!Array.isArray(event)) {
    // DIRECT invoke — handle as single item
    return await resolveOne(event);
  }

  // BATCH invoke — event is an array of contexts
  const serverIds = event.map(ctx => ctx.source.serverId);

  // Batch fetch from DynamoDB using BatchGetItem
  const results = await db.batchGetServerStatuses(serverIds);

  // CRITICAL: return array in SAME ORDER as input
  // If an item failed, return an error object at that position, not null
  return event.map((ctx, i) => {
    const status = results[ctx.source.serverId];
    if (!status) {
      // Return an error at this position — AppSync maps it to GraphQL error for this item
      return {
        data: null,
        errorMessage: `Server ${ctx.source.serverId} not found`,
        errorType: "NotFoundError"
      };
    }
    return {
      data: {
        status: status.isHealthy ? "healthy" : "down",
        lastCheckedAt: status.lastCheckedAt,
        uptimePct: status.uptimePct
      }
    };
  });
  // Each array item must be either:
  // - The plain result object (for success)
  // - { data: result, errorMessage: "...", errorType: "..." } for partial failure
  // Returning null at a position results in a null field (valid for nullable fields)
};

// CDK: enable BATCH on the LambdaDataSource
const ds = new appsync.LambdaDataSource(this, 'StatusDs', {
  api,
  lambdaFunction: statusLambda
});
new appsync.Resolver(this, 'StatusResolver', {
  api,
  typeName: 'Server',
  fieldName: 'status',
  dataSource: ds,
  maxBatchSize: 100  // enables BATCH mode; 0 = DIRECT (default)
});

The order invariant in BATCH mode is absolute: AppSync maps the Lambda's response array back to the original request positions by index. If you return 3 items for a 4-item batch (for example, by filtering out a failed fetch), AppSync will silently assign wrong results to wrong objects. Always return an array of the exact same length as the input, using the { data, errorMessage, errorType } shape for failures.

IAM execution role and service principal

The AppSync Lambda data source requires an IAM service role that AppSync can assume to invoke the Lambda function. The role's trust policy must list appsync.amazonaws.com as the principal. This is separate from the Lambda function's resource-based policy — while you can also add a Lambda resource-based policy granting appsync.amazonaws.com, the service role approach is required for AppSync to call the Lambda across accounts or when the Lambda has no resource-based policy.

// CDK: create service role for AppSync → Lambda invocation
import * as iam from 'aws-cdk-lib/aws-iam';

const appSyncLambdaRole = new iam.Role(this, 'AppSyncLambdaRole', {
  assumedBy: new iam.ServicePrincipal('appsync.amazonaws.com'),
  inlinePolicies: {
    InvokeLambda: new iam.PolicyDocument({
      statements: [
        new iam.PolicyStatement({
          actions: ['lambda:InvokeFunction'],
          resources: [
            toolDispatchLambda.functionArn,
            // Include :* for alias/version invocations if needed
          ]
        })
      ]
    })
  }
});

const lambdaDs = new appsync.LambdaDataSource(this, 'ToolLambdaDs', {
  api,
  lambdaFunction: toolDispatchLambda,
  serviceRole: appSyncLambdaRole  // explicit role; CDK creates one automatically if omitted
});

// If using CDK auto-created role, verify it has lambda:InvokeFunction:
// lambdaDs.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({
//   actions: ['lambda:InvokeFunction'],
//   resources: [toolDispatchLambda.functionArn]
// }));

The 30-second resolver timeout: async tool pattern

AppSync enforces a hard 30-second ceiling on resolver execution, independent of the Lambda's own configured timeout. A tool that takes 45 seconds results in: AppSync returns a resolver timeout error to the caller at 30 seconds, but the Lambda continues running in the background for another 15 seconds. The safe pattern for long-running MCP tools is to return immediately with a job ID, then let the client subscribe for the result.

// Async tool dispatch pattern — return job ID immediately
// 1. Mutation: start job, return jobId
// 2. AppSync subscription: subscribe for onToolComplete(jobId: "...")
// 3. Job Lambda: does the work, writes result, triggers createToolResult mutation

// Mutation resolver (completes in < 1 second)
export const startToolHandler = async (event) => {
  const { toolName, sessionId, toolInput } = event.arguments;

  // Write a pending job record
  const jobId = crypto.randomUUID();
  await db.createToolJob({ jobId, toolName, sessionId, toolInput, status: "pending" });

  // Enqueue to SQS for async processing (not a direct Lambda call)
  await sqs.send(new SendMessageCommand({
    QueueUrl: process.env.TOOL_QUEUE_URL,
    MessageBody: JSON.stringify({ jobId, toolName, toolInput, sessionId })
  }));

  // Return immediately — well within 30 seconds
  return {
    jobId,
    status: "pending",
    createdAt: new Date().toISOString()
  };
};

// Background tool worker Lambda (triggered by SQS, NOT by AppSync)
// This Lambda can run for up to 15 minutes (or configured timeout)
export const toolWorkerHandler = async (event) => {
  const { jobId, toolName, toolInput, sessionId } = JSON.parse(event.Records[0].body);

  const result = await runTool(toolName, toolInput); // may take minutes

  // Write result to AppSync via GraphQL mutation (IAM-signed request)
  // This triggers the subscription: onToolComplete(jobId: jobId)
  await appsyncMutation(`
    mutation CompleteToolJob($input: ToolJobCompleteInput!) {
      completeToolJob(input: $input) { jobId status output durationMs }
    }
  `, { input: { jobId, status: "complete", output: JSON.stringify(result), durationMs: result.durationMs } });
};

Failure modes reference

FailureSymptomFix
Lambda throws unhandled exceptionGraphQL error with errorType: "Lambda:Unhandled" and the exception messageCatch all errors and return structured error objects; throw { message, errorType, data } for typed GraphQL errors
BATCH response array shorter than input arrayWrong results silently mapped to wrong parent objectsAlways return array of exact same length as input; use { data: null, errorMessage, errorType } for missing items
AppSync 30-second timeout — Lambda still runningCaller gets resolver timeout error; Lambda continues in background; duplicate side effects on retryDesign mutations to be idempotent (deduplicate by clientMutationId); use async pattern (job queue + subscription) for operations > 5 seconds
Service role missing lambda:InvokeFunctionAll resolver invocations fail with Access denied at the data source call stageGrant lambda:InvokeFunction on the Lambda ARN to the AppSync service role (appsync.amazonaws.com principal)
Non-nullable field missing in Lambda return valueResolver error: "Cannot return null for non-nullable type"Include all non-nullable fields in the return object; use undefined (not null) for absent optional fields — AppSync omits undefined values from the response
DIRECT event received in BATCH handler (or vice versa)TypeError: event.map is not a functionCheck Array.isArray(event) to distinguish BATCH vs DIRECT events; handle both in the same Lambda or use separate Lambdas per mode
Lambda cold start exceeds 30 seconds under BATCH with large itemsFirst batch after idle period times out; subsequent batches succeedEnable Lambda provisioned concurrency for AppSync data source Lambdas with latency-sensitive resolvers