AWS AppSync · 2026-10-02 · AppSync advanced arc

AWS AppSync for MCP Servers: Pipeline Resolvers, Multi-Mode Authorization, and Production Caching

AppSync is the managed GraphQL layer that can orchestrate your entire MCP tool invocation pipeline — session auth, tool dispatch, and audit write — in a single GraphQL field without any server managing those steps. The platform handles real-time subscriptions natively (so MCP clients can subscribe to tool results), supports five independent authorization modes on individual fields, and serves read-heavy status APIs from an in-memory cache that costs a fraction of the DynamoDB reads it replaces. But AppSync has a narrow contract around resolvers: the JavaScript runtime's helper namespace is the only way to call util.error() correctly, the BATCH Lambda resolver requires returning results in the exact same index order as input, and the AppSync-imposed 30-second ceiling applies independently of your Lambda's own configured timeout. This guide synthesizes five AppSync topics — pipeline resolvers, Lambda resolvers, authorization modes, server-side caching, and conflict detection — into three structural patterns that cover everything from a cold request landing on the pipeline to a cache-invalidating mutation updating shared MCP server config.

TL;DR

Pattern 1 — Pipeline resolver and Lambda resolver architecture

The 3-function canonical MCP pipeline

A pipeline resolver executes a chain of AppSync functions in sequence, each function calling its own data source. For an MCP tool invocation, the canonical chain is: (1) session auth check against DynamoDB, (2) tool dispatch via a Lambda data source, (3) audit write back to DynamoDB. All three functions share the same ctx object, and specifically ctx.stash — a plain JavaScript object that persists for the entire pipeline lifetime.

The before-handler runs first, before any function's request handler. It extracts the arguments from ctx.args and validates them:

// Pipeline before-handler — JavaScript AppSync runtime
export function request(ctx) {
  ctx.stash.sessionId = ctx.args.sessionId;
  ctx.stash.toolName = ctx.args.toolName;
  ctx.stash.callerUserId = ctx.identity?.sub ?? null;

  if (!ctx.args.sessionId || !ctx.args.toolName) {
    util.error("Missing required argument", "ArgumentError");
  }

  return {};  // MUST return {} — any non-empty object is treated as a data source request
}

export function response(ctx) {
  if (ctx.error) {
    util.appendError(ctx.error.message, ctx.error.type, ctx.result);
  }
  return ctx.result;  // final return value: whatever the last function put in ctx.result
}

Function 1 reads the session from DynamoDB and validates ownership. The key distinction here is util.error() versus util.appendError(): util.error() aborts the pipeline immediately and no subsequent functions run. util.appendError() adds a GraphQL error to the response but the pipeline continues to the next function. Use util.error() in auth checks — you never want tool dispatch to proceed if the session is expired or belongs to another user.

// Function 1: session auth check (DynamoDB data source)
export function request(ctx) {
  return {
    operation: "GetItem",
    key: { sessionId: util.dynamodb.toDynamoDB(ctx.stash.sessionId) }
  };
}

export function response(ctx) {
  const session = ctx.result;
  if (!session) {
    util.error("Session not found", "AuthError");   // aborts — no tool dispatch
  }
  if (session.expiresAt < util.time.nowEpochSeconds()) {
    util.error("Session expired", "AuthError");
  }
  if (session.userId !== ctx.stash.callerUserId) {
    util.error("Session belongs to a different user", "AuthError");
  }
  ctx.stash.session = session;
  ctx.stash.teamId = session.teamId;   // propagated to function 2 and 3
  return session;
}

Function 2 dispatches the tool. It reads arguments and stash data from the shared ctx, calls the Lambda data source, and puts the result into ctx.stash.toolResult for the audit function. Using util.appendError() here (rather than util.error()) allows the audit write in Function 3 to still run even when the tool fails — which is exactly what you want for audit completeness.

// Function 2: tool dispatch (Lambda data source)
export function request(ctx) {
  return {
    operation: "Invoke",
    payload: {
      toolName: ctx.stash.toolName,
      sessionId: ctx.stash.sessionId,
      teamId: ctx.stash.teamId,
      toolInput: ctx.args.input,
      callerUserId: ctx.stash.callerUserId
    }
  };
}

export function response(ctx) {
  if (ctx.error) {
    util.appendError(ctx.error.message, "ToolInvocationError");  // continues to audit
    return null;
  }
  ctx.stash.toolResult = ctx.result;
  return ctx.result;
}

// Function 3: audit write (DynamoDB data source)
export function request(ctx) {
  const record = {
    auditId: util.autoId(),
    sessionId: ctx.stash.sessionId,
    teamId: ctx.stash.teamId,
    toolName: ctx.stash.toolName,
    status: ctx.stash.toolResult?.status ?? "error",
    durationMs: ctx.stash.toolResult?.durationMs ?? 0,
    createdAt: util.time.nowISO8601()
  };
  return {
    operation: "PutItem",
    key: { auditId: util.dynamodb.toDynamoDB(record.auditId) },
    attributeValues: util.dynamodb.toMapValues(record)
  };
}

export function response(ctx) {
  return ctx.stash.toolResult;  // return the tool result, NOT ctx.result (the DynamoDB PutItem response)
}

The last function's return value becomes the pipeline result. Function 3's response handler returns ctx.stash.toolResult, not ctx.result — the DynamoDB PutItem response would be meaningless to the GraphQL caller.

JavaScript AppSync runtime: util.* helpers and ES2022 restrictions

The APPSYNC_JS runtime (version 1.0.0) is an ES2022 subset — it supports if/else, loops, array methods, destructuring, template literals, and async/await for pipeline functions. The two notable restrictions are eval and new Function — both are blocked. The util.* namespace provides DynamoDB type conversion, time utilities, ID generation, and error handling:

// Type conversion
util.dynamodb.toDynamoDB("string")    // { S: "string" }
util.dynamodb.toDynamoDB(42)          // { N: "42" }
util.dynamodb.toMapValues({ k: "v" }) // { k: { S: "v" } }
util.dynamodb.toObject(ctx.result)    // AttributeValue map → plain JS object

// Time + IDs
util.time.nowISO8601()                // "2026-10-02T11:00:00.000Z"
util.time.nowEpochSeconds()           // epoch seconds
util.autoId()                         // random UUID v4

// Error handling
util.error("message", "ErrorType")   // throws — aborts current function and propagates
util.appendError("msg", "Type")      // adds to errors array — continues execution
util.warn("message")                  // CloudWatch log only — execution continues

// JSON
util.toJson(obj)    // JSON.stringify
util.parseJson(str) // JSON.parse
util.base64Encode(str)
util.base64Decode(str)

CDK setup for a pipeline resolver

CDK appsync.Resolver accepts a pipelineConfig array of AppsyncFunction objects. The order of functions in the array is the execution order.

import * as appsync from 'aws-cdk-lib/aws-appsync';

const authFn = new appsync.AppsyncFunction(this, 'AuthFn', {
  api,
  dataSource: sessionTableDs,
  name: 'SessionAuthFunction',
  code: appsync.Code.fromAsset('resolvers/auth-function.js'),
  runtime: appsync.FunctionRuntime.JS_1_0_0
});
const toolFn = new appsync.AppsyncFunction(this, 'ToolFn', {
  api,
  dataSource: toolLambdaDs,
  name: 'ToolDispatchFunction',
  code: appsync.Code.fromAsset('resolvers/tool-function.js'),
  runtime: appsync.FunctionRuntime.JS_1_0_0
});
const auditFn = new appsync.AppsyncFunction(this, 'AuditFn', {
  api,
  dataSource: auditTableDs,
  name: 'AuditWriteFunction',
  code: appsync.Code.fromAsset('resolvers/audit-function.js'),
  runtime: appsync.FunctionRuntime.JS_1_0_0
});

new appsync.Resolver(this, 'InvokeToolResolver', {
  api,
  typeName: 'Mutation',
  fieldName: 'invokeTool',
  pipelineConfig: [authFn, toolFn, auditFn],  // execution order
  code: appsync.Code.fromAsset('resolvers/invoke-tool-pipeline.js'),
  runtime: appsync.FunctionRuntime.JS_1_0_0
});

DIRECT vs BATCH Lambda resolvers: the order invariant

For unit resolvers — where one GraphQL field maps to one Lambda invocation — use DIRECT mode. The Lambda receives a single event with the full GraphQL context: arguments, identity, source (the parent object for nested fields), request.headers, and info.selectionSetList. BATCH mode activates when maxBatchSize > 0 on the resolver: AppSync groups up to that many pending invocations into a single array event. The Lambda must return an array of results in the same index order as the input array — not filtered, not reordered, not shortened.

// Lambda handler — handles both DIRECT and BATCH
export const handler = async (event) => {
  // DIRECT: event is a single context object
  // BATCH: event is an array of context objects
  if (!Array.isArray(event)) {
    return await resolveOne(event);
  }

  // BATCH: fetch all server IDs in one DynamoDB BatchGetItem
  const serverIds = event.map(ctx => ctx.source.serverId);
  const results = await db.batchGetServerStatuses(serverIds);

  // CRITICAL: return array in SAME INDEX ORDER as input
  // event[0] result → response[0], event[1] result → response[1], etc.
  return event.map(ctx => {
    const status = results[ctx.source.serverId];
    if (!status) {
      return {
        data: null,
        errorMessage: `Server ${ctx.source.serverId} not found`,
        errorType: "NotFoundError"
      };
    }
    return {
      data: {
        status: status.isHealthy ? "healthy" : "down",
        lastCheckedAt: status.lastCheckedAt,
        uptimePct: status.uptimePct
      }
    };
  });
  // Each position: plain result object (success) or { data, errorMessage, errorType }
};

The order invariant is absolute: if you filter out items (returning a shorter array), AppSync maps response[0] to the wrong parent object — silently. The only safe handling for a missing or errored item is to return an error object at that index position.

The 30-second ceiling and the async tool pattern

AppSync enforces a hard 30-second ceiling on resolver execution regardless of the Lambda's own configured timeout. A tool that takes 45 seconds produces this behavior: AppSync returns a resolver timeout error at the 30-second mark, and the Lambda continues running in the background for another 15 seconds. The caller never receives the result, but the Lambda incurs the full execution cost.

For MCP tools with unpredictable or long execution times, use the async pattern: the AppSync mutation completes in under one second by queuing work to SQS and returning a jobId. A background Lambda worker processes the tool and fires an IAM-signed GraphQL mutation back to AppSync when done. A subscription on the client side delivers the result.

// Mutation resolver — completes in < 1 second
export const startToolHandler = async (event) => {
  const { toolName, sessionId, toolInput } = event.arguments;
  const jobId = crypto.randomUUID();

  await db.createToolJob({ jobId, toolName, sessionId, status: "pending" });
  await sqs.send(new SendMessageCommand({
    QueueUrl: process.env.TOOL_QUEUE_URL,
    MessageBody: JSON.stringify({ jobId, toolName, toolInput, sessionId })
  }));

  return { jobId, status: "pending", createdAt: new Date().toISOString() };
};

// Background worker Lambda (SQS trigger — NOT AppSync)
export const toolWorkerHandler = async (event) => {
  const { jobId, toolName, toolInput, sessionId } = JSON.parse(event.Records[0].body);
  const result = await runTool(toolName, toolInput); // may take minutes

  // Fire an IAM-signed mutation back to AppSync
  await appsyncMutation(`
    mutation CompleteToolJob($input: ToolJobCompleteInput!) {
      completeToolJob(input: $input) { jobId status output durationMs }
    }
  `, { input: { jobId, status: "complete", output: JSON.stringify(result), durationMs: result.durationMs } });
};

The background worker uses IAM SigV4 to call AppSync (the completeToolJob mutation is annotated @aws_iam). The client subscribes to onToolJobComplete(jobId: $jobId) and receives the result when the worker fires the mutation. This pattern decouples tool execution time from the AppSync timeout entirely.

Pattern 2 — Authorization: multi-mode API design

Five auth modes and field-level directives

AppSync supports five authorization modes: API key (anonymous or machine-to-machine with a static key), Cognito User Pools (JWT tokens from your user pool), IAM SigV4 (AWS IAM role-based access for server-to-server calls), Lambda authorizer (custom token validation), and OIDC (tokens from any standards-compliant OIDC provider). An AppSync API has one default auth mode and up to four additional modes. Field-level @aws_* directives control which modes can access each field — a field with no directive is accessible only via the default mode.

type Query {
  # Public status feed — API key (no auth required)
  listMcpServers(limit: Int, nextToken: String): ServerConnection
    @aws_api_key

  # Both public (API key) and authenticated users (Cognito)
  getMcpServerStatus(serverId: ID!): ServerStatus
    @aws_api_key @aws_cognito_user_pools

  # Private monitoring data — Cognito users only
  getMyServerAlerts(serverId: ID!): [Alert]
    @aws_cognito_user_pools

  # Server-to-server gateway calls — IAM SigV4
  getInternalMetrics(teamId: ID!): Metrics
    @aws_iam
}

type Mutation {
  # Authenticated users update config
  claimMcpServer(serverId: ID!, verificationToken: String!): Server
    @aws_cognito_user_pools

  # Internal gateway writes tool results
  recordToolInvocation(input: ToolInvocationInput!): ToolResult
    @aws_iam

  # Custom token validation via Lambda authorizer
  configureWebhookAlert(serverId: ID!, webhookUrl: String!): Alert
    @aws_lambda
}

type Subscription {
  # CRITICAL: subscription auth is INDEPENDENT from linked mutation auth
  # Must declare its own directives — @aws_subscribe does NOT inherit auth
  onServerStatusChange(serverId: ID!): ServerStatus
    @aws_subscribe(mutations: ["updateServerStatus"])
    @aws_api_key @aws_cognito_user_pools
}

The subscription independence is the most commonly missed requirement: @aws_subscribe only declares which mutations trigger the subscription. Authorization for the subscription connection itself is a separate declaration. An onToolComplete subscription linked to a Cognito-protected completeToolJob mutation still requires its own @aws_cognito_user_pools directive, or all subscription connections will return 401.

Lambda authorizer: resolverContext forwarding and TTL caching

The Lambda authorizer receives the raw authorization token, validates it, and returns a structured response. The resolverContext field in the response is forwarded to every resolver called during that request — accessible as ctx.identity.resolverContext. Use this to pass per-request data (team ID, plan tier, allowed endpoint IDs) without a DynamoDB lookup inside each resolver.

// Lambda authorizer handler
export const handler = async (event) => {
  const { authorizationToken } = event;

  if (!authorizationToken?.startsWith('Bearer ')) {
    return { isAuthorized: false };
  }

  let claims;
  try {
    claims = await verifyApiToken(authorizationToken.replace('Bearer ', ''));
  } catch {
    return { isAuthorized: false };
  }

  return {
    isAuthorized: true,
    resolverContext: {
      teamId: claims.teamId,
      userId: claims.sub,
      plan: claims.plan,                        // "free" | "team" | "enterprise"
      privateEndpointIds: claims.endpoints ?? []
    },
    ttlOverride: 300  // cache this authorization for 5 minutes
    // ttlOverride: 0 to disable caching (max 3600s)
  };
};

// Inside any resolver — access resolverContext without a DB lookup:
export function request(ctx) {
  const { teamId, plan } = ctx.identity.resolverContext;
  if (plan === 'free' && ctx.args.private) {
    util.error("Upgrade required", "PlanLimitError");
  }
  // proceed with teamId already available
  ctx.stash.teamId = teamId;
  return {};
}

The ttlOverride: 300 means AppSync caches the authorization response for 5 minutes. Within that window, subsequent requests using the same token skip the Lambda invocation entirely. Set ttlOverride: 0 to disable caching for short-lived tokens. The performance difference is significant: without caching, every GraphQL request incurs an additional Lambda cold start risk.

Cognito User Pools: claims, groups, and ownership enforcement

For Cognito-authenticated requests, the ctx.identity object contains the verified JWT claims at ctx.identity.claims, and the user's Cognito groups at ctx.identity.cognitoGroups. Always use ctx.identity.cognitoGroups for group checks — not ctx.identity.claims["cognito:groups"], which requires URL-decoding the colon and is not reliably present in all JWT implementations.

// Resolver enforcing server ownership with group-based admin override
export function response(ctx) {
  const server = ctx.result;
  if (!server) util.error("Server not found", "NotFoundError");

  const callerId = ctx.identity.sub;
  const isOwner = server.ownerId === callerId;
  const isAdmin = (ctx.identity.cognitoGroups ?? []).includes('admin');

  if (!isOwner && !isAdmin) {
    util.error("Access denied", "AuthorizationError");
  }
  return server;
}

// ctx.identity shape for Cognito auth:
// {
//   sub: "user-uuid",
//   username: "johndoe",
//   claims: { sub, email, "cognito:groups": ["authors", "team-abc"], … },
//   cognitoGroups: ["authors", "team-abc"],   // <-- use this for group checks
//   defaultAuthStrategy: "ALLOW"
// }

IAM SigV4: server-to-server MCP gateway calls

The MCP gateway Lambda (or any AWS service that needs to write tool results to AppSync) uses IAM SigV4 signing. The IAM role policy must grant appsync:GraphQL at the field ARN level — a wildcard on the API ARN works but grants more than needed.

// IAM policy for the MCP gateway role — field-scoped
{
  "Effect": "Allow",
  "Action": "appsync:GraphQL",
  "Resource": [
    "arn:aws:appsync:us-east-1:123456789:apis/API_ID/types/Mutation/fields/recordToolInvocation",
    "arn:aws:appsync:us-east-1:123456789:apis/API_ID/types/Query/fields/getInternalMetrics"
  ]
}

// Signing AppSync requests in Node.js with aws4
import aws4 from 'aws4';
import { fromNodeProviderChain } from '@aws-sdk/credential-providers';

const creds = await fromNodeProviderChain()();
const body = JSON.stringify({ query, variables });

const signed = aws4.sign({
  host: 'XXXXXXXX.appsync-api.us-east-1.amazonaws.com',
  path: '/graphql',
  service: 'appsync',
  region: 'us-east-1',
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body
}, {
  accessKeyId: creds.accessKeyId,
  secretAccessKey: creds.secretAccessKey,
  sessionToken: creds.sessionToken  // required for assumed-role credentials
});

const response = await fetch(`https://${signed.host}${signed.path}`, {
  method: 'POST',
  headers: signed.headers,
  body
});

The sessionToken field is required when the Lambda uses an assumed-role credential (which is always the case for Lambda execution roles). Omitting it produces a 403 IncompleteSignature error even when the role policy is correct.

Pattern 3 — Caching, concurrency control, and conflict resolution

PER_RESOLVER_CACHING: per-field TTL and cache key design

AppSync's FULL_REQUEST caching mode caches the entire response for an operation — the same query string and variables return the cached response for all callers within the TTL. PER_RESOLVER_CACHING gives each resolver field its own TTL and cache key, allowing different fields in the same query to have different expiry windows. Enable the cache at the API level and configure it per-resolver in CDK:

// CDK: enable PER_RESOLVER_CACHING at the API level
new appsync.CfnApiCache(this, 'ApiCache', {
  apiId: api.apiId,
  type: 'T2_SMALL',                       // ~$0.038/hr — suitable for <10K cached entries
  ttl: 300,                               // default TTL
  apiCachingBehavior: 'PER_RESOLVER_CACHING',
  atRestEncryptionEnabled: true,
  transitEncryptionEnabled: true
});

// Public status field — 60-second TTL, cache key = serverId only (no identity)
new appsync.Resolver(this, 'StatusResolver', {
  api,
  typeName: 'Query',
  fieldName: 'getMcpServerStatus',
  dataSource: statusDynamoDs,
  code: appsync.Code.fromAsset('resolvers/get-status.js'),
  runtime: appsync.FunctionRuntime.JS_1_0_0,
  cachingConfig: {
    ttl: cdk.Duration.seconds(60),
    cachingKeys: ['$context.arguments.serverId']  // shared across all callers
  }
});

// Private alerts field — 30-second TTL, cache key includes user identity
new appsync.Resolver(this, 'AlertsResolver', {
  api,
  typeName: 'Query',
  fieldName: 'getMyAlerts',
  dataSource: alertsDynamoDs,
  code: appsync.Code.fromAsset('resolvers/get-alerts.js'),
  runtime: appsync.FunctionRuntime.JS_1_0_0,
  cachingConfig: {
    ttl: cdk.Duration.seconds(30),
    cachingKeys: ['$context.identity.sub', '$context.arguments.userId']
  }
});

The cache key trap is including caller identity in the cache key for public data. If getMcpServerStatus includes $context.identity.sub, every unique caller gets their own cache entry — the cache stores N separate entries for the same server status, each containing identical data. Hit rate drops to near zero. For public data shared across callers, use only the query arguments as cache keys.

Cache invalidation: evictFromApiCache in mutation response handlers

AppSync does not automatically invalidate cache entries when the underlying data changes. After a mutation updates an MCP server config, any getMcpServerStatus queries for that server would return stale data for up to the TTL duration. Fix this by calling extensions.evictFromApiCache() in the mutation's response handler.

// Mutation response handler — evict stale cache entries after write
export function response(ctx) {
  const { serverId } = ctx.args.input;
  const updatedServer = ctx.result;

  // Evict the cached entry for getMcpServerStatus on this server
  extensions.evictFromApiCache('Query', 'getMcpServerStatus', {
    '$context.arguments.serverId': serverId  // MUST exactly match the cachingKeys expression
  });

  // Evict server details cache if separately cached
  extensions.evictFromApiCache('Query', 'getMcpServerDetails', {
    '$context.arguments.serverId': serverId
  });

  return updatedServer;
}

The eviction key expression strings must exactly match the cachingKeys strings in the resolver configuration — character for character, including the $context. prefix. A mismatch causes evictFromApiCache to silently succeed without actually removing the entry.

ElastiCache tier selection and break-even calculation

Instance type Memory ~Cost/hr Suitable for
T2_SMALL 1.5 GB ~$0.038 Dev/staging; <10K entries
T2_MEDIUM 3.22 GB ~$0.077 Small production; <50K entries
R4_LARGE 12.3 GB ~$0.30 Production with large result sets
R4_XLARGE 25.6 GB ~$0.59 High-throughput status APIs

Break-even for T2_SMALL at ~$0.038/hr: if DynamoDB on-demand reads cost $0.25 per million RCUs and you serve 10M status reads per hour, the cache achieves break-even if the hit rate exceeds 0.3%. In practice, MCP server status APIs hit rates of 80–95% for hot servers — the cache pays for itself within the first few minutes of traffic. Monitor with CloudWatch CachingHits and CachingMisses metrics; alarm if the hit rate drops below 50% for three consecutive 5-minute windows.

// CDK: CloudWatch hit rate alarm
const cacheHits = new cloudwatch.Metric({
  namespace: 'AWS/AppSync',
  metricName: 'CachingHits',
  dimensionsMap: { GraphQLAPIId: api.apiId },
  statistic: 'Sum', period: cdk.Duration.minutes(5)
});
const cacheMisses = new cloudwatch.Metric({
  namespace: 'AWS/AppSync',
  metricName: 'CachingMisses',
  dimensionsMap: { GraphQLAPIId: api.apiId },
  statistic: 'Sum', period: cdk.Duration.minutes(5)
});

const hitRate = new cloudwatch.MathExpression({
  expression: 'hits / (hits + misses) * 100',
  usingMetrics: { hits: cacheHits, misses: cacheMisses },
  period: cdk.Duration.minutes(5)
});

new cloudwatch.Alarm(this, 'LowCacheHitRate', {
  metric: hitRate,
  threshold: 50,
  comparisonOperator: cloudwatch.ComparisonOperator.LESS_THAN_THRESHOLD,
  evaluationPeriods: 3,
  treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING
});

Conflict detection: _version lifecycle and OPTIMISTIC_CONCURRENCY

When conflict detection is enabled on a DynamoDB data source, AppSync manages three metadata fields on every item: _version (monotonically incrementing integer), _lastChangedAt (epoch milliseconds), and _deleted (boolean for soft-deletes). Every mutation input must include the current _version value read from a prior query. AppSync checks whether the DynamoDB item's current _version matches the client's submitted _version; if not, a conflict has occurred.

With OPTIMISTIC_CONCURRENCY, AppSync rejects conflicting mutations with a ConflictUnhandled error. The client must re-read the item to get the current _version, apply its changes on top of the latest state, and retry. For MCP server config updates, a 3-retry loop with brief backoff handles the majority of simultaneous edit scenarios:

// Client-side retry for OPTIMISTIC_CONCURRENCY
async function updateAlertConfig(serverId, changes) {
  for (let retries = 0; retries < 3; retries++) {
    // Re-read on each attempt to get current _version
    const current = await appsyncQuery(GET_CONFIG_QUERY, { serverId });
    const currentVersion = current.data.getServerConfig._version;

    try {
      return await appsyncMutation(UPDATE_CONFIG_MUTATION, {
        input: { serverId, ...changes, _version: currentVersion }
      });
    } catch (err) {
      if (err.errors?.[0]?.errorType === 'ConflictUnhandled') {
        await sleep(100 * (retries + 1));  // 100ms, 200ms, 300ms
        continue;
      }
      throw err;
    }
  }
  throw new Error('Failed to update after 3 retries — too many concurrent writers');
}

LAMBDA conflict handler: field-level merge for multi-subsystem config

When multiple subsystems update different fields of the same MCP server config record (for example, one service updates alertWebhookUrl while another updates slackChannel), OPTIMISTIC_CONCURRENCY causes unnecessary rejections — both writes are semantically valid because they touch different fields. A Lambda conflict handler receives both the client's proposed item (newItem) and the current DynamoDB item (existingItem), and returns a merged result.

// Lambda conflict handler — field-level last-writer-wins merge
export const handler = async (event) => {
  // event.typeName: "McpServerConfig"
  // event.fieldName: "updateServerConfig"
  // event.operation: "UPDATE"
  // event.newItem: what the client tried to write (_version is the client's stale version)
  // event.existingItem: current DynamoDB item (_version is the actual current version)

  const { newItem, existingItem } = event;

  const merged = {
    ...existingItem,
    // Only overwrite fields that the client explicitly set (non-null in newItem)
    ...(newItem.alertWebhookUrl != null && { alertWebhookUrl: newItem.alertWebhookUrl }),
    ...(newItem.alertThresholdMs != null && { alertThresholdMs: newItem.alertThresholdMs }),
    ...(newItem.slackChannel != null && { slackChannel: newItem.slackChannel }),
    // DO NOT set _version — AppSync manages this
  };

  return merged;   // AppSync writes merged item and increments _version
  // return null to reject (equivalent to OPTIMISTIC_CONCURRENCY behavior)
};

// CDK: attach Lambda conflict handler to DynamoDB data source
const cfnDs = configTableDs.node.defaultChild as appsync.CfnDataSource;
cfnDs.dynamoDbConfig = {
  tableName: serverConfigTable.tableName,
  awsRegion: this.region,
  conflictDetection: 'VERSION',
  conflictHandler: 'LAMBDA',
  lambdaConflictHandlerArn: conflictLambda.functionArn
};

_deleted items and TTL cleanup

AppSync soft-deletes items by setting _deleted: true rather than removing them from DynamoDB. For hand-written MCP schemas (not using Amplify DataStore), these deleted items accumulate permanently unless you add a DynamoDB TTL attribute. Filter them from list queries explicitly:

// DynamoDB scan/query filter for list resolvers
// FilterExpression: "attribute_not_exists(#d) OR #d = :false"
// ExpressionAttributeNames: { "#d": "_deleted" }
// ExpressionAttributeValues: { ":false": { BOOL: false } }

// In a JS resolver response handler:
export function response(ctx) {
  // Single item: filter soft-deleted
  if (ctx.result?._deleted) return null;
  return ctx.result;
}

Add DynamoDB TTL on deleted items by setting a _ttl attribute (epoch seconds) in the conflict handler or a DynamoDB Streams processor: if newItem._deleted === true, set _ttl = Math.floor(Date.now() / 1000) + 7 * 86400. DynamoDB will remove the item within 48 hours of the TTL expiry.

Consolidated failure modes

Symptom Root cause Fix
Pipeline resolver before-handler CloudFormation error Before-handler returns non-empty object — AppSync treats it as a data source request document Return {} from the before-handler's request function
Pipeline aborts when tool fails — no audit record written Function 2 calls util.error() on tool failure Use util.appendError() in Function 2's response — pipeline continues to Function 3
Function 3 returns DynamoDB AttributeValue map instead of tool result Response handler returns ctx.result (PutItem response) instead of ctx.stash.toolResult Return ctx.stash.toolResult from Function 3's response handler
BATCH resolver returns wrong data for some parent objects Lambda response array shorter than input array or items returned in wrong order Map over event array by index — never filter, reorder, or shorten the response array
Caller receives timeout error at 30 seconds; Lambda still running AppSync 30-second ceiling independent of Lambda timeout Use async pattern: mutation queues work to SQS, subscription delivers result when worker completes
All subscription connections return 401 Subscription field missing its own auth directive — @aws_subscribe does not inherit auth Add @aws_cognito_user_pools (or other mode) directly on the subscription field
Lambda authorizer invoked on every request despite TTL ttlOverride: 0 or authorizer not returning a ttlOverride field Return ttlOverride: 300 from authorizer for 5-minute cache window
Cognito group check always false Using ctx.identity.claims["cognito:groups"] instead of ctx.identity.cognitoGroups Use ctx.identity.cognitoGroups ?? [] for group membership checks
IAM SigV4 returns 403 IncompleteSignature sessionToken omitted from signing credentials on assumed-role invocation Pass sessionToken: creds.sessionToken to aws4.sign()
Cache entries not shared across users for public data $context.identity.sub included in cachingKeys for public fields Exclude identity from cache keys on public data — use only query arguments
Stale cache after mutation Mutation resolver not calling extensions.evictFromApiCache() Add eviction call in mutation response handler with exact-match key expressions
evictFromApiCache succeeds but stale entry persists Key expression string mismatch between eviction call and cachingKeys config Copy-paste exact strings from CDK cachingKeys into eviction call
Conflict detection silently disabled for some mutations Mutation input missing _version field Always include _version in mutation input types; add it as a required field in the GraphQL schema
Lambda conflict handler returns rejection error Handler returns undefined instead of null Return null explicitly to reject; return the merged item to accept
Deleted server configs appear in list queries _deleted: true items not filtered — AppSync never removes soft-deleted items Add FilterExpression: "attribute_not_exists(#d) OR #d = :false" on list queries; set DynamoDB TTL on _deleted items

Production checklists

Pipeline resolver checklist

Lambda resolver checklist

Authorization checklist

Caching and conflict detection checklist

How AliveMCP fits in

AliveMCP monitors the protocol-layer health of your MCP servers — not the HTTP surface that AppSync's built-in health check probes, but the full JSON-RPC lifecycle: initialize handshake, tools/list schema response, and synthetic tool execution. An AppSync API that returns 200 on every request can still have every tool call silently returning malformed results, expired session tokens, or schema drift between your resolver and the actual tool Lambda. AliveMCP catches the difference. See plans →