Guide · AWS AppSync · Server-Side Caching
AppSync Server-Side Caching for MCP Servers — FULL_REQUEST, Per-Resolver, and Cache Key Design
AppSync server-side caching stores resolver results in an ElastiCache cluster managed by AppSync, returning cached responses for repeated queries without invoking the underlying data source. For an MCP server status API — where many clients poll the same getMcpServerStatus(serverId: "popular-server") query every few seconds — caching collapses that fan-out into a single DynamoDB read per TTL window. There are three caching levels: FULL_REQUEST (the entire GraphQL response is cached by query + variables + identity; one cache entry per unique query string/variables/caller), PER_RESOLVER_CACHING (each resolver field can declare its own TTL and custom cache keys; different fields in the same query can have different TTLs), and NONE (caching disabled). The critical operational rule: a cache entry is NOT automatically invalidated when the underlying data changes. You must explicitly flush affected cache entries after a mutation using extensions.evictFromApiCache() in the mutation resolver — otherwise clients receive stale data for the entire TTL duration.
TL;DR
Enable caching at the AppSync API level with a TTL and instance type (T2_SMALL starts at ~$0.038/hr). Use FULL_REQUEST for simple read-heavy APIs; use PER_RESOLVER_CACHING for mixed read/write APIs where mutations should not invalidate unrelated cached queries. Cache keys default to the query string + variables + caller identity — for public status reads, exclude the identity from the cache key so all callers share the same cache entry. Call extensions.evictFromApiCache(typeName, fieldName, cacheKeys) in mutation resolvers to invalidate affected entries. Monitor 4XXError and cache hit rate via CloudWatch CachingHits and CachingMisses metrics.
Caching levels: FULL_REQUEST vs PER_RESOLVER_CACHING
FULL_REQUEST caches the complete GraphQL response for the entire operation. If the same query string and variables are sent again within the TTL, AppSync returns the cached response without running any resolvers. This is the simplest setup but has a major drawback: a mutation or subscription sent in the same operation as a query bypasses the cache — and the cache is keyed on the full normalized query string, so adding a whitespace or field alias breaks the cache key.
# Schema: caching TTL and cache key directives
# PER_RESOLVER_CACHING is declared at the field level in the resolver config
type Query {
# Short TTL: MCP server status changes every 60 seconds
getMcpServerStatus(serverId: ID!): ServerStatus
# Longer TTL: server metadata changes rarely
getMcpServerDetails(serverId: ID!): ServerDetails
# No cache: live data, always fresh
getMyAlerts(userId: ID!): [Alert]
}
# Resolver-level caching is configured in the resolver definition (CDK/CloudFormation),
# not in the schema. The schema just defines field types.
// CDK: enable PER_RESOLVER_CACHING with custom TTL per field
const api = new appsync.GraphqlApi(this, 'McpApi', {
// ... other config
xrayEnabled: true
});
// Enable caching at the API level (required for any resolver caching)
const cfnApi = api.node.defaultChild as appsync.CfnGraphQLApi;
// Caching is configured via the CfnApiCache resource
new appsync.CfnApiCache(this, 'ApiCache', {
apiId: api.apiId,
type: 'T2_SMALL', // instance type — see cost table below
ttl: 300, // default TTL in seconds (overridable per resolver)
apiCachingBehavior: 'PER_RESOLVER_CACHING', // or FULL_REQUEST
atRestEncryptionEnabled: true,
transitEncryptionEnabled: true
});
// Resolver with caching enabled and custom cache key
new appsync.Resolver(this, 'StatusResolver', {
api,
typeName: 'Query',
fieldName: 'getMcpServerStatus',
dataSource: statusDynamoDs,
code: appsync.Code.fromAsset('resolvers/get-status.js'),
runtime: appsync.FunctionRuntime.JS_1_0_0,
cachingConfig: {
ttl: cdk.Duration.seconds(60),
// Cache key: only serverId (not caller identity — public data shared across callers)
cachingKeys: ['$context.arguments.serverId']
}
});
// Private resolver: cache key includes user identity
new appsync.Resolver(this, 'AlertsResolver', {
api,
typeName: 'Query',
fieldName: 'getMyAlerts',
dataSource: alertsDynamoDs,
code: appsync.Code.fromAsset('resolvers/get-alerts.js'),
runtime: appsync.FunctionRuntime.JS_1_0_0,
cachingConfig: {
ttl: cdk.Duration.seconds(30),
// Cache per user: include identity in key
cachingKeys: ['$context.identity.sub', '$context.arguments.userId']
}
});
Cache key expressions use the same $context (abbreviated $ctx) variable path syntax as VTL mapping templates. Common cache key components: $context.arguments.FIELD, $context.identity.sub, $context.identity.claims.teamId, $context.source.serverId (for nested resolver caching). If cachingKeys is empty, the cache key defaults to the full request hash (query + variables + identity).
Flush-on-mutation: invalidating cache entries after writes
AppSync does not automatically invalidate cache entries when data changes. A mutation that updates a server's status must explicitly evict the cached getMcpServerStatus entry for that server, or callers will receive stale data for the TTL duration. Use extensions.evictFromApiCache() in the mutation resolver's response handler.
// Mutation resolver response handler — evict cached queries after write
// Called after the DynamoDB PutItem/UpdateItem completes
export function response(ctx) {
const { serverId } = ctx.args.input;
const updatedServer = ctx.result;
// Evict the cached getMcpServerStatus entry for this server
// This is a fire-and-forget — AppSync evicts asynchronously
extensions.evictFromApiCache('Query', 'getMcpServerStatus', {
'$context.arguments.serverId': serverId
});
// If the server details are also cached, evict those too
extensions.evictFromApiCache('Query', 'getMcpServerDetails', {
'$context.arguments.serverId': serverId
});
// Return the mutation result to the caller
return updatedServer;
}
// evictFromApiCache signature:
// extensions.evictFromApiCache(
// typeName: string, // e.g., "Query"
// fieldName: string, // e.g., "getMcpServerStatus"
// cacheKeys: object // key-value pairs matching the resolver's cachingKeys config
// )
// The cacheKeys object keys must exactly match the cachingKeys expressions in the resolver config
A common mistake: the cache key expression strings in evictFromApiCache must exactly match those in the resolver's cachingKeys configuration. If the resolver uses $context.arguments.serverId as the cache key, the eviction call must use { '$context.arguments.serverId': value } as the cache key object — not { serverId: value }.
ElastiCache instance types and cost
AppSync manages the ElastiCache cluster — you choose the instance type and AppSync provisions and manages it. The instance type determines both the cache size (how many entries fit) and the throughput (requests per second before eviction pressure). AppSync caching is billed per hour regardless of hit rate.
| Instance type | Memory | Approx cost/hr (us-east-1) | Suitable for |
|---|---|---|---|
| T2_SMALL | 1.5 GB | ~$0.038 | Dev/staging; APIs with < 10K entries |
| T2_MEDIUM | 3.22 GB | ~$0.077 | Small production; < 50K entries |
| R4_LARGE | 12.3 GB | ~$0.30 | Production APIs with large result sets |
| R4_XLARGE | 25.6 GB | ~$0.59 | High-throughput status APIs with large schemas |
For an MCP status API serving millions of server status reads, T2_SMALL with a 60-second TTL can handle high traffic since the cache entry size for a typical server status object is small (< 1 KB). The break-even calculation: if DynamoDB on-demand reads cost $0.25/million read request units and you serve 10M requests/hr, a T2_SMALL at $0.038/hr achieves break-even if the cache hit rate exceeds 0.3% — practically always worth it for repeated status queries.
Monitoring cache hit rate with CloudWatch
AppSync publishes caching metrics to CloudWatch under the AWS/AppSync namespace. The key metrics are CachingHits and CachingMisses, both as count metrics. An alarm on low hit rate identifies queries that are not benefiting from caching (because TTL is too short, cache keys are too specific, or the mutation eviction pattern is overly aggressive).
// CDK: CloudWatch alarm for low AppSync cache hit rate
import * as cloudwatch from 'aws-cdk-lib/aws-cloudwatch';
const cacheHits = new cloudwatch.Metric({
namespace: 'AWS/AppSync',
metricName: 'CachingHits',
dimensionsMap: { GraphQLAPIId: api.apiId },
statistic: 'Sum',
period: cdk.Duration.minutes(5)
});
const cacheMisses = new cloudwatch.Metric({
namespace: 'AWS/AppSync',
metricName: 'CachingMisses',
dimensionsMap: { GraphQLAPIId: api.apiId },
statistic: 'Sum',
period: cdk.Duration.minutes(5)
});
// Hit rate math expression
const hitRate = new cloudwatch.MathExpression({
expression: 'hits / (hits + misses) * 100',
usingMetrics: { hits: cacheHits, misses: cacheMisses },
period: cdk.Duration.minutes(5)
});
new cloudwatch.Alarm(this, 'LowCacheHitRate', {
metric: hitRate,
threshold: 50,
comparisonOperator: cloudwatch.ComparisonOperator.LESS_THAN_THRESHOLD,
evaluationPeriods: 3,
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING
});
Failure modes reference
| Failure | Symptom | Fix |
|---|---|---|
| Mutation not evicting cache → stale reads | Callers see old server status for up to TTL seconds after update | Call extensions.evictFromApiCache() in every mutation resolver that modifies cached fields; match eviction keys to resolver cachingKeys exactly |
evictFromApiCache key string mismatch | Eviction call silently succeeds but cache entry remains; stale reads continue | The key string must be the exact $context.* expression used in the resolver's cachingKeys array, not the variable name |
| FULL_REQUEST cache with different query formatting | Two identical semantic queries return different cache behavior because whitespace or field order differs | AppSync normalizes query strings for FULL_REQUEST cache keys — test cache behavior with the exact query format your client sends; alias-renamed fields produce different cache keys |
Caching enabled but resolver has no cachingConfig | Under PER_RESOLVER_CACHING mode, resolver is not cached (correct behavior) | In PER_RESOLVER_CACHING mode, only resolvers with explicit cachingConfig are cached — add cachingConfig to resolvers you want cached |
| Identity included in cache key for public data | Each unique caller has their own cache entry; no cache sharing; hit rate near 0% | Exclude $context.identity.* from cache keys for public read-only fields; only include identity in cache keys for user-specific data |
| Cache instance too small — eviction under high traffic | Hit rate drops suddenly; DynamoDB read spikes during cache thrashing | Check ElasticacheEvictions CloudWatch metric; upgrade instance type or reduce TTL to limit cache entry count |