Guide · AWS MediaConvert · CloudWatch Monitoring
AWS MediaConvert CloudWatch Monitoring — Job Metrics, Failure Alerts, and EventBridge Status Events
AWS MediaConvert publishes job metrics to CloudWatch automatically — no agent installation, no log streaming configuration required — but the metrics namespace, available dimensions, and event schema are non-obvious enough that most teams miss the most actionable signals. For MCP server video pipelines, the critical monitoring primitives are: CloudWatch AWS/MediaConvert namespace metrics for queue depth and error rates, EventBridge MediaConvert Job State Change events for per-job status transitions (the right way to detect completion or failure without polling), and cost monitoring via TranscodedBillableMinutes to catch runaway transcoding before the bill arrives. MediaConvert does not write to CloudWatch Logs directly — job-level error details come from EventBridge events, not log groups. The monitoring pattern is: EventBridge → Lambda → DynamoDB for operational state, plus CloudWatch alarms for aggregate signals like error rate spike or queue depth growth that indicate systemic pipeline problems rather than individual job failures.
TL;DR
MediaConvert metrics are in the AWS/MediaConvert namespace with dimension Queue set to the queue ARN. Key metrics: JobsCompletedCount, JobsErroredCount, TranscodedBillableMinutes, HDOutputDuration, SDOutputDuration. Set a CloudWatch alarm on JobsErroredCount with threshold 1 and period 300 seconds for a 5-minute error alert. Use EventBridge rule matching source: aws.mediaconvert + detail-type: MediaConvert Job State Change for per-job completion handling — not polling. Include detail.userMetadata in your job creation call to tag jobs with application context that flows through to EventBridge events. See the S3 workflow guide for the Lambda that creates jobs and handles completion events.
CloudWatch metrics in AWS/MediaConvert namespace
MediaConvert publishes six standard metrics to the AWS/MediaConvert CloudWatch namespace. The Queue dimension is the queue ARN — metrics are per-queue, not per-account or per-job. This means if you use multiple queues (e.g., high-priority and batch), you get separate metric streams per queue and can alert independently on each.
# List available metrics in the MediaConvert namespace
aws cloudwatch list-metrics \
--namespace "AWS/MediaConvert" \
--query 'Metrics[*].{MetricName:MetricName,Dimensions:Dimensions}' \
--output table
# Key metrics:
# JobsCompletedCount — Number of COMPLETE jobs per period
# JobsErroredCount — Number of ERROR jobs per period
# TranscodedBillableMinutes — Billable output minutes (HD + SD combined)
# HDOutputDuration — Output minutes >= 720p (billed at HD rate)
# SDOutputDuration — Output minutes < 720p (billed at SD rate)
# StandbyTime — Reserved queue idle time (relevant cost for reserved queues)
# Get job completion rate for the default queue over last hour
QUEUE_ARN="arn:aws:mediaconvert:us-east-1:123456789012:queues/Default"
aws cloudwatch get-metric-statistics \
--namespace "AWS/MediaConvert" \
--metric-name "JobsCompletedCount" \
--dimensions "Name=Queue,Value=$QUEUE_ARN" \
--start-time "$(date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%SZ)" \
--end-time "$(date -u +%Y-%m-%dT%H:%M:%SZ)" \
--period 300 \
--statistics Sum \
--query 'Datapoints[*].{Time:Timestamp,Completed:Sum}' \
--output table
# Get error count over last hour
aws cloudwatch get-metric-statistics \
--namespace "AWS/MediaConvert" \
--metric-name "JobsErroredCount" \
--dimensions "Name=Queue,Value=$QUEUE_ARN" \
--start-time "$(date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%SZ)" \
--end-time "$(date -u +%Y-%m-%dT%H:%M:%SZ)" \
--period 300 \
--statistics Sum
Metrics are reported with a 1-minute granularity but AWS recommends using 5-minute (300-second) periods for alarm evaluation to reduce noise from single-job errors. Metrics for a queue appear only when jobs have been processed — a queue with no activity in the last hour has no datapoints for that period. Set TreatMissingData: notBreaching on alarms so silent queues do not trigger alerts.
Creating CloudWatch alarms for error rate and cost
Two alarms cover the majority of MediaConvert pipeline operational concerns: an error rate alarm that fires when any job fails (useful for small pipelines where every error matters) and a cost alarm that fires when billable minutes exceed your expected budget for the period (useful for catching runaway jobs or unexpected input files that produce very long outputs).
QUEUE_ARN="arn:aws:mediaconvert:us-east-1:123456789012:queues/Default"
ALARM_TOPIC="arn:aws:sns:us-east-1:123456789012:ops-alerts"
# Alarm: any job error in the last 5 minutes
aws cloudwatch put-metric-alarm \
--alarm-name "mediaconvert-job-errors" \
--alarm-description "MediaConvert job failed — check EventBridge error events for details" \
--namespace "AWS/MediaConvert" \
--metric-name "JobsErroredCount" \
--dimensions "Name=Queue,Value=$QUEUE_ARN" \
--statistic Sum \
--period 300 \
--evaluation-periods 1 \
--threshold 1 \
--comparison-operator GreaterThanOrEqualToThreshold \
--treat-missing-data notBreaching \
--alarm-actions "$ALARM_TOPIC"
# Alarm: billable minutes exceeding cost threshold (e.g., > 60 mins in 5 minutes)
# This catches runaway jobs processing unexpectedly long input files
aws cloudwatch put-metric-alarm \
--alarm-name "mediaconvert-cost-spike" \
--alarm-description "MediaConvert billable minutes exceeded 60min in 5min period — possible runaway transcoding" \
--namespace "AWS/MediaConvert" \
--metric-name "TranscodedBillableMinutes" \
--dimensions "Name=Queue,Value=$QUEUE_ARN" \
--statistic Sum \
--period 300 \
--evaluation-periods 1 \
--threshold 60 \
--comparison-operator GreaterThanThreshold \
--treat-missing-data notBreaching \
--alarm-actions "$ALARM_TOPIC"
# Alarm: queue depth (jobs in SUBMITTED state) growing (requires custom metric)
# MediaConvert does not publish queue depth natively — use a scheduled Lambda
# that calls list-jobs --status SUBMITTED and puts a custom metric
aws cloudwatch put-metric-alarm \
--alarm-name "mediaconvert-queue-backlog" \
--namespace "Custom/MediaConvert" \
--metric-name "SubmittedJobCount" \
--dimensions "Name=Queue,Value=Default" \
--statistic Maximum \
--period 300 \
--evaluation-periods 3 \
--threshold 20 \
--comparison-operator GreaterThanThreshold \
--treat-missing-data notBreaching \
--alarm-actions "$ALARM_TOPIC"
For cost-sensitive MCP pipelines, add input duration validation before calling create-job: use ffprobe or the AWS Rekognition Video API to check the input video duration. If the input exceeds your maximum expected duration (e.g., 30 minutes), reject it at the MCP tool level before a job is created — this prevents users from submitting 4-hour raw recordings that trigger large transcoding charges.
EventBridge job state change events
MediaConvert emits state change events to EventBridge for every job lifecycle transition: SUBMITTED, PROGRESSING, COMPLETE, ERROR, and CANCELED. The event payload includes the full job details — output file paths, duration, error code, and any user metadata you attached at job creation. This is the authoritative source for job status — use it instead of polling get-job.
# EventBridge event structure for MediaConvert Job State Change:
{
"version": "0",
"id": "event-id",
"source": "aws.mediaconvert",
"account": "123456789012",
"time": "2026-10-03T10:00:00Z",
"region": "us-east-1",
"detail-type": "MediaConvert Job State Change",
"detail": {
"timestamp": 1696300000000,
"accountId": "123456789012",
"queue": "arn:aws:mediaconvert:us-east-1:123456789012:queues/Default",
"jobId": "1696300000000-abc123",
"status": "COMPLETE",
"userMetadata": {
"userId": "user-456",
"uploadId": "upload-789",
"correlationId": "req-abc"
},
"outputGroupDetails": [{
"type": "FILE_GROUP",
"outputDetails": [{
"outputFilePaths": ["s3://my-video-output/processed/video-001/_1080p.mp4"],
"durationInMs": 120000,
"videoDetails": {
"widthInPx": 1920,
"heightInPx": 1080,
"averageBitrate": 4850000
}
}]
}]
}
}
# For ERROR events, detail includes:
# "status": "ERROR",
# "errorCode": 1040,
# "errorMessage": "Job encountered an error..."
# For PROGRESSING events, detail includes:
# "status": "PROGRESSING",
# "jobProgress": {
# "jobPercentComplete": 45,
# "currentPhase": "TRANSCODING",
# "phaseProgress": { "TRANSCODING": { "status": "PROGRESSING", "percentComplete": 45 } }
# }
The userMetadata object — which you set at job creation time — flows through to every EventBridge event for that job. Use it to correlate EventBridge events back to your application's data model without a DynamoDB lookup. For example, include userId and uploadId in the metadata so your completion Lambda can update the correct database record and notify the correct user without needing to join through a job ID.
Attaching user metadata to jobs for correlation
The UserMetadata parameter in create-job accepts up to 10 key-value pairs. Each key and value is a string with a 256-character limit. This metadata is stored with the job and included in all EventBridge events and in the get-job response — it is the primary mechanism for correlating MediaConvert jobs with your application's records without an external database lookup.
# Attach application metadata at job creation time
aws mediaconvert create-job \
--endpoint-url "$ENDPOINT" \
--role "$ROLE_ARN" \
--user-metadata '{
"userId": "user-456",
"uploadId": "upload-789",
"contentType": "tutorial",
"requestedAt": "2026-10-03T10:00:00Z",
"callbackUrl": "https://api.example.com/webhooks/video-processed"
}' \
--settings '...'
# Completion Lambda that reads userMetadata from EventBridge event:
def handle_completion(event):
detail = event["detail"]
metadata = detail.get("userMetadata", {})
user_id = metadata.get("userId")
upload_id = metadata.get("uploadId")
callback_url = metadata.get("callbackUrl")
output_paths = []
for group in detail.get("outputGroupDetails", []):
for output in group.get("outputDetails", []):
output_paths.extend(output.get("outputFilePaths", []))
# Update database
update_video_status(upload_id, "COMPLETE", output_paths)
# Notify user
if callback_url:
requests.post(callback_url, json={
"uploadId": upload_id,
"status": "COMPLETE",
"outputs": output_paths
})
return {"statusCode": 200}
Include a callbackUrl in the metadata if your MCP tool creates jobs on behalf of external callers that need webhook-style notification. The completion Lambda can POST to this URL directly, decoupling the notification delivery from the job-creation Lambda and avoiding long-polling by the caller. Validate the callback URL at job creation time (check it is an allowed domain) to prevent SSRF via injected URLs in user-controlled metadata.
Failure modes reference
| Failure | Symptom | Fix |
|---|---|---|
| No metrics in AWS/MediaConvert namespace | CloudWatch shows no data for JobsCompletedCount | Metrics only appear after a job has been processed in the queue — submit a test job; also verify the Queue dimension value is the full ARN, not just the queue name |
| Alarm stuck in INSUFFICIENT_DATA | Error rate alarm never transitions to OK or ALARM | No jobs have run recently — set TreatMissingData: notBreaching so the alarm goes to OK when no metrics are reported; also verify the queue dimension ARN matches exactly (the ARN is case-sensitive and region-specific) |
| EventBridge rule not matching events | MediaConvert jobs complete but completion Lambda never fires | EventBridge rule event pattern must match exactly: "source": ["aws.mediaconvert"] (as array not string), "detail-type": ["MediaConvert Job State Change"]; also check the rule is in the same region as the MediaConvert jobs |
| userMetadata missing from EventBridge events | completion Lambda cannot find userId or uploadId in event detail | userMetadata must be passed in create-job as a flat string-to-string map — nested objects are not supported; also check that the key names match exactly (case-sensitive) |
| PROGRESSING event not firing | EventBridge only receives SUBMITTED and COMPLETE, never PROGRESSING | PROGRESSING events fire when the job begins actual transcoding — for short videos (<30s), the job may complete before the first PROGRESSING event is emitted; this is normal and does not indicate a problem |
| Cost alarm firing on legitimate workloads | TranscodedBillableMinutes alarm triggers every hour | Adjust alarm threshold to match expected peak workload — if you regularly process 60+ minutes of video per 5-minute window, increase the threshold or switch to a daily budget metric; alternatively, use an anomaly detection alarm to catch unexpected spikes relative to baseline |