AWS GuardDuty · 2026-10-08 · GuardDuty arc

AWS GuardDuty for MCP Server Security: Detector Setup, Runtime Protection, and Automated Remediation

The single most dangerous default in a GuardDuty setup is the SIX_HOURS finding-publishing frequency — which means a compromised MCP server running on EC2 can exfiltrate IAM role credentials for six full hours before you see a finding in the console or receive an EventBridge event, giving an attacker more than enough time to pivot across your entire AWS organization via the stolen role. Change FindingPublishingFrequency to FIFTEEN_MINUTES when enabling the detector — this is the single highest-impact configuration change in the entire GuardDuty setup. GuardDuty is a regional threat-detection service that analyzes CloudTrail management events, VPC flow logs, DNS query logs, and S3 data-plane events without requiring any agent installation or log forwarding — it operates independently of your CloudWatch Logs setup and uses its own internal data delivery path. For MCP server operators, GuardDuty covers five distinct security surfaces: (1) the base detector for EC2, IAM, and general AWS API threats — detector setup and finding export (CreateDetector with FIFTEEN_MINUTES frequency, data source pricing table, finding type naming convention ThreatPurpose:ResourceType/ThreatFamilyName, severity levels 7.0–8.9 High / 4.0–6.9 Medium / 0.1–3.9 Low, CreateFilter ARCHIVE suppression for monitoring probes, S3 export with KMS CMK and bucket policy for guardduty.amazonaws.com, Organizations delegated administrator pattern); (2) the findings API — querying, filtering, and deduplication (ListFindings filter criteria for severity/type/service.archived/updatedAt/accountId/resource, GetFindings max 50 per call, deduplication behavior — same finding type + resource + source = update not new finding = same findingId fires EventBridge every 15 minutes for ongoing threats, service.count as true incident volume, idempotency via DynamoDB conditional put, ArchiveFindings/UpdateFindingsFeedback patterns, finding type table for InstanceCredentialExfiltration/CryptoCurrency/Backdoor/Recon/Policy); (3) Kubernetes-level protection — EKS audit log monitoring and runtime DaemonSet agent (two independent layers: audit log monitoring = Kubernetes control-plane API analysis covering RBAC attacks and privileged container deployment, runtime monitoring = DaemonSet agent per node monitoring process execution and network connections inside running containers, enable via UpdateDetector features EKS_AUDIT_LOGS + RUNTIME_MONITORING + EKS_ADDON_MANAGEMENT, Fargate limitation: no DaemonSet support so only audit log findings cover Fargate pods, agent coverage verification via DaemonSet pod count vs node count, audit log finding types AdminAccessToDefaultServiceAccount/AnonymousAccessGranted/PrivilegedContainer/ExecIntoContainer, runtime finding types NewBinaryExecutedInContainer/C&CActivity.B/ProcessInjected/CryptoCurrencyMining, suppression filters for CI/CD kubectl exec and container image scanners); (4) serverless threat detection — Lambda network activity monitoring (no agent or code change required, account-wide coverage, cost model $0/month first 1M invocations then $0.20/1M above, finding types CryptoCurrency:Lambda/BitcoinTool = code injection indicator / Backdoor:Lambda/C&CActivity.B / UnauthorizedAccess:Lambda/MetadataDNSRebind = SSRF via DNS rebinding / TorClientError / SuspiciousNetworkActivity.B = behavioral false-positive during 2-week baseline, false positive handling for LLM API outbound connections by suppressing SuspiciousNetworkActivity.B by function name prefix or ASN but never suppressing CryptoCurrency/Backdoor); and (5) automated response — EventBridge-driven remediation with idempotent Lambda (EventBridge rule source=aws.guardduty detail-type=GuardDuty Finding severity>=7, finding event includes full detail so no GetFindings call needed from Lambda, fires on updates too so idempotency via DynamoDB conditional put is mandatory, IAM credential revocation via PutRolePolicy Deny all with DateLessThan aws:TokenIssueTime = effective within 1–2 seconds and reversible, EC2 isolation via ModifyInstanceAttribute replacing all SGs with pre-created all-deny group not terminate = preserves forensic evidence with SSM still accessible via VPC endpoint, multi-account assume-role with ExternalId condition, deploy remediation role via CloudFormation StackSets). This guide synthesizes all five topics into three structural patterns: detector setup and finding management, runtime protection for EKS and Lambda workloads, and EventBridge-driven automated remediation with idempotent response patterns.

TL;DR

Pattern 1 — Detector Setup and Finding Management

The FIFTEEN_MINUTES frequency requirement and why SIX_HOURS is dangerous

GuardDuty's default finding-publishing frequency is SIX_HOURS. This means that when GuardDuty detects an active threat — an EC2 instance's IAM role credentials being used from outside AWS, a Lambda function making DNS queries to a cryptocurrency mining pool, a container establishing a reverse shell — you will not receive an EventBridge event or see the finding in the console for up to six hours after initial detection. For MCP servers, which handle authenticated user sessions and hold AWS IAM role credentials in their execution environment, a six-hour window before credential revocation is long enough for an attacker to exfiltrate all accessible data, spin up infrastructure in other regions, and establish persistent access.

# Enable GuardDuty with FIFTEEN_MINUTES finding frequency — this is the critical flag
aws guardduty create-detector \
  --enable \
  --finding-publishing-frequency FIFTEEN_MINUTES \
  --data-sources '{
    "S3Logs": {"Enable": true},
    "Kubernetes": {"AuditLogs": {"Enable": true}},
    "MalwareProtection": {"ScanEc2InstanceWithFindings": {"EbsVolumes": {"Enable": true}}}
  }' \
  --tags "Environment=production,Service=mcp-platform"

# If detector already exists with SIX_HOURS, update it
aws guardduty update-detector \
  --detector-id abc1234567890abcdef1234567890 \
  --finding-publishing-frequency FIFTEEN_MINUTES

# Verify the frequency
aws guardduty get-detector \
  --detector-id abc1234567890abcdef1234567890 \
  --query 'FindingPublishingFrequency'
# Expected: "FIFTEEN_MINUTES"

GuardDuty is regional — a detector in us-east-1 does not cover resources in eu-west-1. Enable a detector in every region where your MCP infrastructure runs. The detector ID returned by CreateDetector is used in all subsequent API calls in that region; save it in your infrastructure configuration (SSM Parameter Store or environment variables for your operations Lambda) rather than hardcoding it.

Data source selection: enabling only what generates findings

GuardDuty's three core data sources — CloudTrail management events, VPC flow logs, and DNS query logs — are included in the base detector price. Additional data sources (S3 data-plane events, Kubernetes audit logs, Lambda network activity, EBS malware scanning) each add cost. The correct approach is to enable only the data sources that match your actual MCP infrastructure, not to enable everything by default.

Data sourceIncluded in baseEnable for MCP if
CloudTrail management eventsYesAlways — covers IAM API calls, resource changes, unauthorized API activity
VPC flow logsYesAlways — covers network-based threats for ECS/EC2 MCP deployments
DNS query logsYesAlways — covers C2 communication and crypto mining detection via DNS
S3 data-plane eventsNo (free 30 days)If MCP server accesses S3 buckets with sensitive data or user uploads
Kubernetes audit logs (EKS)NoIf running MCP servers on EKS — covers RBAC attacks and privileged container deployment
Lambda network activityNoIf using Lambda-based MCP functions — covers code injection and SSRF via DNS rebinding
EBS malware scanningNoIf running stateful MCP servers on EC2 with EBS volumes containing persistent data

Pre-emptive suppression filters: set before first scan cycle

AliveMCP's core function — probing MCP server endpoints — generates network traffic that GuardDuty flags as reconnaissance: Recon:EC2/PortProbeUnprotectedPort and Recon:EC2/Portscan for every scan cycle. Without suppression rules in place before your first scan, you will be flooded with false positives that drown out real threats.

The important distinction between CreateFilter with ARCHIVE action and retroactive archiving: filters run at finding generation time. A finding matching an ARCHIVE filter is immediately suppressed and never appears in the active findings list or fires an EventBridge event. Retroactive archiving only hides findings that have already appeared and already fired EventBridge events — your on-call rotation has already been paged. Always use CreateFilter for known false-positive patterns, not retroactive archiving.

# Create suppression filter for monitoring probe activity — do this BEFORE first scan
aws guardduty create-filter \
  --detector-id abc1234567890abcdef1234567890 \
  --name "monitoring-probe-suppression" \
  --action ARCHIVE \
  --rank 1 \
  --description "Suppress port probe findings from AliveMCP monitoring infrastructure" \
  --finding-criteria '{
    "Criterion": {
      "type": {
        "Equals": ["Recon:EC2/PortProbeUnprotectedPort", "Recon:EC2/Portscan"]
      },
      "resource.instanceDetails.tags.key": {"Equals": ["Service"]},
      "resource.instanceDetails.tags.value": {"Equals": ["mcp-monitoring"]}
    }
  }'

# Suppress sample findings generated by CreateSampleFindings (for testing)
aws guardduty create-filter \
  --detector-id abc1234567890abcdef1234567890 \
  --name "sample-findings-suppression" \
  --action ARCHIVE \
  --rank 2 \
  --finding-criteria '{
    "Criterion": {
      "type": {"Equals": ["SampleFinding"]}
    }
  }'

Querying findings programmatically and deduplication behavior

GuardDuty deduplicates findings: the same threat type targeting the same resource from the same source generates one finding, not multiple. Subsequent occurrences increment service.count and update service.eventLastSeen. A finding with count: 847 represents 847 individual events — the finding ID is constant across all of them. The service.isNew field distinguishes first detection from update: true means this is the first time this finding appeared in the current publishing cycle; false means it is an update.

This deduplication behavior has a critical implication for EventBridge-triggered automation: the same finding ID fires an EventBridge event on every update cycle (every 15 minutes if the threat is ongoing). Your remediation Lambda will be invoked repeatedly for the same incident. Idempotency is not optional — see Pattern 3 for the DynamoDB conditional put pattern.

# List High-severity active findings, newest first
aws guardduty list-findings \
  --detector-id abc1234567890abcdef1234567890 \
  --finding-criteria '{
    "Criterion": {
      "severity": {"Gte": 7},
      "service.archived": {"Eq": ["false"]}
    }
  }' \
  --sort-criteria '{"AttributeName": "updatedAt", "OrderBy": "DESC"}' \
  --max-results 50

# Fetch details for specific IDs (max 50 per call)
aws guardduty get-findings \
  --detector-id abc1234567890abcdef1234567890 \
  --finding-ids "abc123" "def456" \
  --query 'Findings[*].{
    Id:Id,
    Type:Type,
    Severity:Severity,
    Count:Service.Count,
    LastSeen:Service.EventLastSeen,
    IsNew:Service.IsNew
  }'

# Check if ongoing: eventLastSeen within the last 15 minutes = still active
# Count alone does not tell you if the threat is still happening
aws guardduty get-findings \
  --detector-id abc1234567890abcdef1234567890 \
  --finding-ids "abc123" \
  --query 'Findings[0].Service.EventLastSeen'

The most critical finding types for MCP server infrastructure, mapped to immediate actions:

Finding typeSeverityMCP implicationImmediate action
InstanceCredentialExfiltration:EC2/NoInstanceProfileHigh (8.0)EC2 IAM role credentials used from outside AWS — leaked from MCP env vars, code, or logsRevoke all sessions via IAM deny policy (see Pattern 3); rotate secret; quarantine instance
CryptoCurrency:Lambda/BitcoinToolHigh (7.5)Lambda MCP function calling mining pool DNS — code injection via user-controlled tool argDisable function; inspect handler for eval()/exec() with user-controlled args
Backdoor:EC2/C&CActivity.B!DNSHigh (8.0)MCP server host querying known C2 domains — malware installed on hostIsolate instance via SG replacement (see Pattern 3); preserve for forensics
Policy:IAMUser/RootCredentialUsageHigh (8.0)Root credentials used — should never happen in productionReview CloudTrail for root activity; lock with MFA; investigate what root did
Recon:EC2/PortProbeUnprotectedPortMedium (2.0–5.0)Scanning bot or targeted reconnaissance on open MCP server portsReview SGs for unnecessary open ports; add suppression if expected monitoring activity

Pattern 2 — Runtime Protection for EKS and Lambda Workloads

EKS protection: two independent layers that cover different threat surfaces

GuardDuty EKS protection is two entirely separate features, not a single toggle. EKS Audit Log Monitoring analyzes the Kubernetes control-plane API stream — who created which pods, RBAC changes, privileged container deployments, kubectl exec operations. EKS Runtime Monitoring installs a DaemonSet security agent on each worker node that monitors in-container behavior at runtime: process execution, library loading, file system access, and network connections from within running containers. Enabling one does not enable the other. Both are needed for comprehensive MCP server protection on EKS: audit logs catch the initial compromise vector (malicious pod deployment, RBAC escalation), while runtime monitoring catches post-exploitation activity inside the container (download and execution of attack tools, C2 communication, crypto mining).

# Enable EKS Audit Log Monitoring
aws guardduty update-detector \
  --detector-id abc1234567890abcdef1234567890 \
  --features '[{"Name": "EKS_AUDIT_LOGS", "Status": "ENABLED"}]'

# Enable EKS Runtime Monitoring with managed add-on (GuardDuty manages agent lifecycle)
# EKS_ADDON_MANAGEMENT: ENABLED = GuardDuty auto-deploys aws-guardduty-agent DaemonSet
aws guardduty update-detector \
  --detector-id abc1234567890abcdef1234567890 \
  --features '[
    {
      "Name": "RUNTIME_MONITORING",
      "Status": "ENABLED",
      "AdditionalConfiguration": [
        {"Name": "EKS_ADDON_MANAGEMENT", "Status": "ENABLED"}
      ]
    }
  ]'

# Verify the DaemonSet is running on all nodes (run after enabling)
# kubectl get daemonset aws-guardduty-agent --namespace amazon-guardduty -o wide
# Pod count must equal worker node count — silent missing coverage if not equal

# Check add-on is installed on the cluster
aws eks list-addons --cluster-name mcp-production-cluster \
  --query 'addons[?contains(@, `aws-guardduty`)]'

The Fargate limitation is important: GuardDuty runtime monitoring requires a DaemonSet, and Fargate does not support DaemonSets. Fargate-based MCP server pods are covered only by EKS Audit Log Monitoring — there are no runtime findings for Fargate. If your MCP deployment uses a mix of EC2 nodes (for the main service) and Fargate (for batch jobs or auxiliary workloads), understand that runtime coverage is incomplete for the Fargate portion.

EKS audit log finding types and suppression for CI/CD

Audit log findings represent threats at the Kubernetes control plane level — before any container code runs. The most immediately actionable finding types for MCP server operators:

Finding typeWhat triggered itMCP server implication
Policy:Kubernetes/AdminAccessToDefaultServiceAccountClusterRoleBinding grants cluster-admin to default service accountAny pod in that namespace can make arbitrary Kubernetes API calls — escape path for compromised MCP container
PrivilegeEscalation:Kubernetes/PrivilegedContainerContainer deployed with securityContext.privileged: trueContainer has full root access to host — MCP server deployments should never require privileged mode
Execution:Kubernetes/ExecIntoContainerkubectl exec or API equivalent into running containerLegitimate for debugging by operators; suppress for your CI/CD principal; investigate if principal is unexpected
Persistence:Kubernetes/ContainerWithSensitiveMountContainer with hostPath mount to sensitive paths (/etc/passwd, /proc)Container can read/write host credentials — check if MCP server deployment uses any hostPath mounts

Execution:Kubernetes/ExecIntoContainer is one of the highest-volume legitimate-but-flagged findings in production EKS environments. CI/CD pipelines commonly run kubectl exec for post-deployment health checks. Create a suppression filter for your CI/CD IAM role before the first deployment to prevent the finding from polluting your security dashboard:

# Suppress ExecIntoContainer for CI/CD principal (GitHub Actions IAM role)
aws guardduty create-filter \
  --detector-id abc1234567890abcdef1234567890 \
  --name "cicd-kubectl-exec-suppression" \
  --action ARCHIVE \
  --rank 1 \
  --finding-criteria '{
    "Criterion": {
      "type": {"Equals": ["Execution:Kubernetes/ExecIntoContainer"]},
      "resource.kubernetesDetails.kubernetesUserDetails.username": {
        "Prefix": ["arn:aws:sts::123456789012:assumed-role/github-actions-role/"]
      }
    }
  }'

# Suppress NewBinaryExecutedInContainer for image scanners (Trivy, Snyk)
# Scanners may execute binaries during scan — filter by scanner pod name prefix
aws guardduty create-filter \
  --detector-id abc1234567890abcdef1234567890 \
  --name "image-scanner-suppression" \
  --action ARCHIVE \
  --rank 2 \
  --finding-criteria '{
    "Criterion": {
      "type": {"Equals": ["Execution:Kubernetes/NewBinaryExecutedInContainer"]},
      "resource.kubernetesDetails.kubernetesWorkloadDetails.name": {
        "Prefix": ["trivy-scan-", "snyk-scan-"]
      }
    }
  }'

# Suppress PrivilegedContainer for infrastructure pods (kube-system namespace)
# vpc-cni, node-local-dns, kube-proxy require privileged mode in some configs
aws guardduty create-filter \
  --detector-id abc1234567890abcdef1234567890 \
  --name "infra-privileged-suppression" \
  --action ARCHIVE \
  --rank 3 \
  --finding-criteria '{
    "Criterion": {
      "type": {"Equals": ["PrivilegeEscalation:Kubernetes/PrivilegedContainer"]},
      "resource.kubernetesDetails.kubernetesWorkloadDetails.namespace": {
        "Equals": ["kube-system"]
      }
    }
  }'

Lambda protection: enabling without agent installation and managing the baseline period

GuardDuty Lambda protection monitors the network activity of all Lambda functions in an account by analyzing Lambda's internal network flow data — not VPC flow logs, which means it covers both VPC-attached and public Lambda functions without any configuration on the Lambda side. No Lambda layer, no code change, no cold start penalty. The protection operates entirely below the function execution level and has zero impact on MCP server response latency.

# Enable Lambda network activity monitoring — account-wide, no per-function selection
aws guardduty update-detector \
  --detector-id abc1234567890abcdef1234567890 \
  --features '[{"Name": "LAMBDA_NETWORK_LOGS", "Status": "ENABLED"}]'

# Cost: first 1M invocations/month free, $0.20/1M above
# For an MCP server at 100 req/day (3,000/month): entirely free
# For a high-traffic MCP at 1M+ req/day: ~$18-$180/month

# For Organizations: enable Lambda protection on all member accounts
aws guardduty update-organization-configuration \
  --detector-id <admin-detector-id> \
  --features '[{"Name": "LAMBDA_NETWORK_LOGS", "AutoEnable": "ALL"}]'

The two-week baseline period is the primary source of false positives after enabling Lambda protection. GuardDuty's behavioral detection model (Execution:Lambda/SuspiciousNetworkActivity.B) generates findings for outbound connections to IP ranges not seen in the first two weeks of monitoring. An MCP Lambda making calls to LLM provider APIs (api.anthropic.com via Cloudflare, Amazon Bedrock at AWS endpoints) will trigger these findings during the baseline period. The solution is targeted suppression filters by ASN for known-good providers:

# Suppress behavioral false positives for known LLM API providers during baseline
# Cloudflare (AS13335) hosts api.anthropic.com; AWS AS16509 for Bedrock; AS14618 for AWS
aws guardduty create-filter \
  --detector-id abc1234567890abcdef1234567890 \
  --name "mcp-ai-provider-asn-suppression" \
  --action ARCHIVE \
  --rank 1 \
  --finding-criteria '{
    "Criterion": {
      "type": {"Equals": ["Execution:Lambda/SuspiciousNetworkActivity.B"]},
      "service.action.networkConnectionAction.remoteIpDetails.organization.asn": {
        "Equals": ["13335", "16509", "14618"]
      }
    }
  }'

# After the baseline period (2-3 weeks), SuspiciousNetworkActivity.B findings
# naturally decrease as GuardDuty learns your Lambda's normal outbound patterns.
# NEVER suppress CryptoCurrency:Lambda/BitcoinTool or Backdoor:Lambda/C&CActivity.B —
# these are threat-intelligence-based (not behavioral) and have very low false-positive rates.

The critical Lambda finding types that indicate real compromise rather than baseline noise:

Finding typeTriggerMCP context and response
CryptoCurrency:Lambda/BitcoinToolDNS queries to cryptocurrency mining pool domainsHigh-confidence code injection — MCP tool handler passing user input to exec()/subprocess(); disable function, inspect for eval() with user-controlled args
Backdoor:Lambda/C&CActivity.BNetwork connections to known C2 server IPs/domainsFunction compromised via injected dependency or supply-chain attack; disable and audit recent npm/pip package changes
UnauthorizedAccess:Lambda/MetadataDNSRebindDNS queries resolving to 169.254.169.254 (EC2 metadata endpoint)SSRF via DNS rebinding — attacker targeting metadata service; inspect URL-fetch tools that accept user-controlled URLs
UnauthorizedAccess:Lambda/TorClientErrorConnections to Tor entry guard nodesAttacker anonymizing data exfiltration via Tor; block outbound if Lambda is VPC-attached; audit injection vectors

Pattern 3 — EventBridge Automated Remediation with Idempotent Lambda

EventBridge rule structure and why idempotency is mandatory

GuardDuty publishes every finding to the default EventBridge event bus in the same region as the detector, with source: "aws.guardduty" and detail-type: "GuardDuty Finding". The full finding JSON is included in the event's detail field — the remediation Lambda receives everything it needs without calling GetFindings.

The pattern fires on two occasions: when a finding is first generated, and when an existing finding is updated (every 15 minutes if the threat is ongoing, triggered by the FindingPublishingFrequency cycle). A sustained brute-force attack against an MCP server that runs for two hours generates eight EventBridge events for the same finding ID — the Lambda is invoked eight times. Without idempotency, the isolation security group is applied, then when the next event fires the Lambda checks the instance, sees it is isolated, and re-applies the same action — or worse, the idempotency check is wrong and the Lambda un-isolates and re-isolates the instance eight times while the attack is ongoing.

# EventBridge rule for High-severity GuardDuty findings (severity >= 7)
aws events put-rule \
  --name "guardduty-high-severity-findings" \
  --event-pattern '{
    "source": ["aws.guardduty"],
    "detail-type": ["GuardDuty Finding"],
    "detail": {
      "severity": [{"numeric": [">=", 7]}]
    }
  }' \
  --state ENABLED

# Route to both SNS (human alert) and Lambda (auto-remediation)
aws events put-targets \
  --rule "guardduty-high-severity-findings" \
  --targets '[
    {
      "Id": "sns-alert",
      "Arn": "arn:aws:sns:us-east-1:123456789012:security-alerts"
    },
    {
      "Id": "lambda-remediation",
      "Arn": "arn:aws:lambda:us-east-1:123456789012:function:guardduty-auto-remediation"
    }
  ]'

# Separate rule for credential-exfiltration findings at any severity — always auto-remediate
aws events put-rule \
  --name "guardduty-credential-exfiltration" \
  --event-pattern '{
    "source": ["aws.guardduty"],
    "detail-type": ["GuardDuty Finding"],
    "detail": {
      "type": [
        "InstanceCredentialExfiltration:EC2/NoInstanceProfile",
        "InstanceCredentialExfiltration:EC2/ScheduledEvent",
        "UnauthorizedAccess:IAMUser/InstanceCredentialExfiltration"
      ]
    }
  }' \
  --state ENABLED

DynamoDB idempotency check and the remediation Lambda structure

The idempotency mechanism is a DynamoDB table with findingId as the partition key, using a conditional put with attribute_not_exists(findingId). The first invocation for a finding ID succeeds and proceeds with remediation. Every subsequent invocation for the same finding ID throws ConditionalCheckFailedException — the Lambda logs this and returns without action. The DynamoDB item TTL should be 90 days (matching GuardDuty's finding retention period) so the table does not grow unbounded.

// Node.js Lambda — idempotent GuardDuty auto-remediation
const { DynamoDBClient, PutItemCommand } = require('@aws-sdk/client-dynamodb');
const { EC2Client, ModifyInstanceAttributeCommand, DescribeInstanceAttributeCommand } = require('@aws-sdk/client-ec2');
const { IAMClient, PutRolePolicyCommand } = require('@aws-sdk/client-iam');
const { SNSClient, PublishCommand } = require('@aws-sdk/client-sns');

const dynamo = new DynamoDBClient({});
const ec2 = new EC2Client({});
const iam = new IAMClient({});
const sns = new SNSClient({});

exports.handler = async (event) => {
  const finding = event.detail;
  const findingId = finding.id;
  const findingType = finding.type;
  const severity = finding.severity;
  const accountId = finding.accountId;

  // Idempotency check: conditional put fails on duplicate findingId
  try {
    await dynamo.send(new PutItemCommand({
      TableName: process.env.RESPONSES_TABLE,
      Item: {
        findingId: { S: findingId },
        findingType: { S: findingType },
        severity: { N: String(severity) },
        respondedAt: { S: new Date().toISOString() },
        accountId: { S: accountId },
        ttl: { N: String(Math.floor(Date.now() / 1000) + 90 * 24 * 3600) }
      },
      ConditionExpression: 'attribute_not_exists(findingId)'
    }));
  } catch (e) {
    if (e.name === 'ConditionalCheckFailedException') {
      console.log(`Already responded to finding ${findingId} — skipping`);
      return;  // Not an error — expected for ongoing threats
    }
    throw e;
  }

  // Route to remediation action based on finding type
  if (findingType.includes('InstanceCredentialExfiltration') ||
      findingType.includes('UnauthorizedAccess:IAMUser/InstanceCredentialExfiltration')) {
    await revokeIAMCredentials(finding);
  } else if (findingType.includes('Backdoor:EC2') || findingType.includes('C&CActivity')) {
    await isolateEC2Instance(finding);
  } else if (findingType.includes('CryptoCurrency:Lambda') || findingType.includes('Backdoor:Lambda')) {
    await disableLambdaFunction(finding);
  }

  // Always send SNS notification with action taken
  await sns.send(new PublishCommand({
    TopicArn: process.env.ALERT_TOPIC_ARN,
    Subject: `GuardDuty Auto-Remediation: ${findingType}`,
    Message: JSON.stringify({ findingId, findingType, severity, accountId,
      action: 'auto-remediated', timestamp: new Date().toISOString() }, null, 2)
  }));
};

IAM credential revocation: the PutRolePolicy deny pattern

When InstanceCredentialExfiltration:EC2/NoInstanceProfile fires, the EC2 instance's IAM role credentials are being used outside of AWS. The fastest mitigation is an IAM inline deny policy with an explicit Deny that overrides all Allow statements. The key condition is DateLessThan aws:TokenIssueTime: <now> — this denies all actions performed by STS tokens that were issued before the timestamp, effectively invalidating all active sessions for that role within 1–2 seconds.

A common point of confusion about this condition: it denies credentials issued BEFORE the timestamp. New credentials issued AFTER the policy is applied have a newer issue time and are NOT denied by the condition — the MCP server continues to function normally on new sessions. This is why the pattern is superior to deleting the IAM role: the role remains intact and new legitimate sessions work, while all leaked sessions (which have an older issue time) are immediately blocked. Removing the inline policy restores full access when the investigation completes.

async function revokeIAMCredentials(finding) {
  const roleName = extractRoleName(finding.resource?.instanceDetails);
  if (!roleName) {
    console.error('Could not extract role name from finding', finding.id);
    return;
  }

  // Deny all actions for all sessions issued before NOW
  // New sessions (issued after this policy is applied) are NOT affected
  const denyPolicy = {
    Version: '2012-10-17',
    Statement: [{
      Effect: 'Deny',
      Action: '*',
      Resource: '*',
      Condition: {
        DateLessThan: {
          'aws:TokenIssueTime': new Date().toISOString()
        }
      }
    }]
  };

  await iam.send(new PutRolePolicyCommand({
    RoleName: roleName,
    PolicyName: `GuardDutyRevoke-${finding.id.substring(0, 8)}`,
    PolicyDocument: JSON.stringify(denyPolicy)
  }));

  console.log(`Revoked all sessions for role ${roleName} (finding: ${finding.id})`);
  // To restore: aws iam delete-role-policy --role-name <role> --policy-name GuardDutyRevoke-<id>
}

function extractRoleName(instanceDetails) {
  const iamProfile = instanceDetails?.iamInstanceProfile;
  if (!iamProfile?.arn) return null;
  const match = iamProfile.arn.match(/role\/(.+)$/);
  return match ? match[1] : null;
}

EC2 instance isolation: security group replacement over termination

For backdoor and C2 findings on EC2-based MCP servers, the correct response is network isolation via security group replacement — not termination. Terminating the instance destroys the memory state, process table, and ephemeral file system that contain forensic evidence of how the attacker got in, what they accessed, and what they did. Security group replacement cuts all inbound and outbound network access while keeping the instance running for forensic investigation.

The isolation security group is a security group with no inbound rules and no outbound rules. AWS EC2 security groups have an implicit deny for traffic not matched by any rule — a group with zero rules denies all traffic in both directions. Create this group once per VPC before you need it, not during incident response when time pressure causes mistakes.

# Create the isolation security group once per VPC — do this during setup, not during incident
aws ec2 create-security-group \
  --group-name "guardduty-isolation" \
  --description "GuardDuty isolation group — no inbound or outbound rules — do not add rules" \
  --vpc-id vpc-abc123 \
  --tag-specifications 'ResourceType=security-group,Tags=[{Key=Purpose,Value=GuardDutyIsolation}]'

# Save the returned GroupId in SSM Parameter Store for the remediation Lambda
aws ssm put-parameter \
  --name "/security/guardduty-isolation-sg-id" \
  --value "sg-isolation-id" \
  --type String

# Verify no rules exist on the isolation group
aws ec2 describe-security-groups \
  --group-ids sg-isolation-id \
  --query 'SecurityGroups[0].{Inbound:IpPermissions,Outbound:IpPermissionsEgress}'
# Expected: {Inbound: [], Outbound: []}
async function isolateEC2Instance(finding) {
  const instanceId = finding.resource?.instanceDetails?.instanceId;
  if (!instanceId) return;

  const isolationSgId = process.env.ISOLATION_SECURITY_GROUP_ID;

  // Check if already isolated (API-level idempotency — DynamoDB check already passed,
  // but this handles edge cases where the SG was already replaced by another system)
  const { Groups } = await ec2.send(new DescribeInstanceAttributeCommand({
    InstanceId: instanceId,
    Attribute: 'groupSet'
  }));

  if (Groups.some(g => g.GroupId === isolationSgId)) {
    console.log(`Instance ${instanceId} already isolated`);
    return;
  }

  // Save original SGs in DynamoDB for later restoration
  await dynamo.send(new PutItemCommand({
    TableName: process.env.RESPONSES_TABLE,
    Item: {
      findingId: { S: `isolation-${instanceId}` },
      originalSecurityGroups: { S: JSON.stringify(Groups.map(g => g.GroupId)) },
      instanceId: { S: instanceId },
      isolatedAt: { S: new Date().toISOString() }
    }
  }));

  // Replace all SGs with isolation group — instance loses all normal network access
  await ec2.send(new ModifyInstanceAttributeCommand({
    InstanceId: instanceId,
    Groups: [isolationSgId]
  }));

  console.log(`Isolated ${instanceId} — replaced ${Groups.length} SG(s) with isolation group`);
  // Forensic access: SSM Session Manager still works via SSM VPC endpoint (no SG rules needed)
  // To restore: aws ec2 modify-instance-attribute --instance-id <id> --groups <original-ids>
}

The SSM access preservation point is important: AWS Systems Manager Session Manager connects via the SSM endpoint URL (not directly to port 22 on the instance). If your VPC has SSM, EC2Messages, and SSMMessages VPC endpoints configured, the SSM agent on the isolated instance can still initiate an outbound HTTPS connection to the SSM endpoint — which does not require open security group rules. Without VPC endpoints, SSM traffic goes through the internet-facing SSM endpoint via NAT gateway, which is blocked by the isolation group. Create the SSM VPC endpoints before relying on this forensic access path.

Multi-account remediation via cross-account role assumption

With a GuardDuty delegated administrator setup, all member account findings appear in the admin account's EventBridge. The finding's accountId field identifies which member account the affected resource belongs to. The remediation Lambda runs in the admin account and must assume a cross-account role in the member account to execute EC2 or IAM actions there.

// Multi-account remediation: assume GuardDutyRemediationRole in affected account
const { STSClient, AssumeRoleCommand } = require('@aws-sdk/client-sts');
const sts = new STSClient({});

async function getClientForAccount(accountId, region, ClientClass) {
  const roleArn = `arn:aws:iam::${accountId}:role/GuardDutyRemediationRole`;
  const externalId = process.env.REMEDIATION_EXTERNAL_ID;  // Prevents confused-deputy attacks

  const { Credentials } = await sts.send(new AssumeRoleCommand({
    RoleArn: roleArn,
    RoleSessionName: `guardduty-remediation-${Date.now()}`,
    ExternalId: externalId,
    DurationSeconds: 900  // 15 minutes — enough for one remediation action
  }));

  return new ClientClass({
    region,
    credentials: {
      accessKeyId: Credentials.AccessKeyId,
      secretAccessKey: Credentials.SecretAccessKey,
      sessionToken: Credentials.SessionToken
    }
  });
}

// Usage:
// const memberEC2 = await getClientForAccount(finding.accountId, finding.region, EC2Client);
// await memberEC2.send(new ModifyInstanceAttributeCommand({ ... }));

// Required: GuardDutyRemediationRole in every member account with trust policy:
// Principal: { "AWS": "arn:aws:iam::SECURITY-ACCOUNT:role/guardduty-auto-remediation" }
// Condition: { "StringEquals": { "sts:ExternalId": "SHARED-SECRET" } }
// Permissions: ec2:ModifyInstanceAttribute, ec2:DescribeInstanceAttribute,
//              iam:PutRolePolicy, lambda:UpdateFunctionConfiguration

Deploy GuardDutyRemediationRole to all member accounts via CloudFormation StackSets from the management account — not manually. The ExternalId condition prevents confused-deputy attacks where a third party tricks your Lambda into assuming the role for an incident in a different account. Both the trust policy and the Lambda environment variable must use the same ExternalId value.

Consolidated failure modes reference

FailureSymptomRoot cause and fix
Findings not appearing for 6 hours after compromiseActive credential exfiltration with no GuardDuty alertDetector using default SIX_HOURS publishing frequency; run update-detector --finding-publishing-frequency FIFTEEN_MINUTES immediately
S3 findings export stuck at PENDING_VERIFICATIONPublishing destination never reaches PUBLISHING stateKMS key policy missing GuardDuty service principal for GenerateDataKey/Decrypt, or S3 bucket policy missing PutObject for guardduty.amazonaws.com; check both
Alert flood from monitoring probes every scan cycleDozens of Recon:EC2/PortProbeUnprotectedPort per hourNo suppression filter before first scan; create CreateFilter ARCHIVE rule matching monitoring instance tags immediately
Remediation Lambda triggered 8 times for one 2-hour attackSame isolation action applied repeatedlyNo DynamoDB idempotency check; add conditional put with attribute_not_exists(findingId) before any remediation action
EKS runtime monitoring enabled but no in-container findingsKnown container events not generating runtime findingsDaemonSet pod count does not equal node count — agent missing on some nodes provides no coverage silently; verify via kubectl get daemonset in amazon-guardduty namespace
Lambda protection finding flood after enablingHundreds of SuspiciousNetworkActivity.B findings on first dayExpected during 2-week baseline learning period; create ASN-based suppression filters for known LLM API providers (AS13335, AS16509, AS14618)
IAM deny policy not blocking compromised credentialsStolen credentials still working after PutRolePolicyVerify policy was applied via get-role-policy; check that DateLessThan condition uses current timestamp (not a hardcoded past date); propagation is typically under 2 seconds globally
Isolated EC2 instance not accessible via SSM after SG replacementCannot forensically investigate instance via Session ManagerSSM requires HTTPS outbound to SSM endpoint — blocked if VPC has no NAT gateway and no SSM/EC2Messages/SSMMessages VPC endpoints; create VPC endpoints before relying on SSM forensic access path
Cross-account assume-role fails during multi-account remediationAccessDenied when Lambda targets member account resourceGuardDutyRemediationRole not created in member account via StackSets, or trust policy references wrong security account ID; verify with aws sts get-caller-identity from remediation Lambda
Member account findings not visible in admin accountAdmin detector shows only its own account findingsMember account not enrolled via Organizations auto-enable; check list-members for DISABLED status; verify update-organization-configuration --auto-enable ALL was applied
Fargate MCP pods generating no runtime findingsEKS runtime monitoring enabled but Fargate pods have no coverageExpected — GuardDuty runtime agent requires DaemonSet which Fargate does not support; Fargate pods are protected only by EKS Audit Log Monitoring findings
GetFindings returns empty array for known finding IDsListFindings returned IDs but GetFindings emptyFindings expire after 90 days and return empty (not an error); verify you are querying the correct detector ID and region