AWS MediaConvert · 2026-10-03 · MediaConvert arc
AWS MediaConvert for MCP Servers: Video Transcoding Pipelines, HLS Streaming, and Production Monitoring
MediaConvert has one surprising property that catches every first-time user: every API call — job creation, queue listing, preset management — must go to an account-specific endpoint URL, not the generic regional endpoint, and calling the generic endpoint returns HTTP 404 Not Found with no indication you hit the wrong host. For MCP server video pipelines, MediaConvert solves the "transcode user-uploaded video without running FFmpeg" problem — a Lambda triggered by S3 calls describe-endpoints once (caching the result across warm invocations), then create-job with an input S3 URI and an output destination, and EventBridge notifies a completion Lambda when the job finishes — no polling, no long-running processes, scales to zero when idle. The five topics that produce a production-ready MediaConvert pipeline — job setup and IAM roles (account-specific endpoint, trust policy for mediaconvert.amazonaws.com, settings JSON structure, queue types), job templates and presets (system preset library, custom H.264 QVBR presets, positional merge for per-job destination override, update-job-template versioning), S3 trigger workflow (Lambda job creator, DynamoDB ETag idempotency guard, iam:PassRole requirement, error code classification), HLS adaptive streaming output (HLS_GROUP_SETTINGS, multi-rendition ladder, GopClosedCadence: 1 for segment alignment, CloudFront cache split), and CloudWatch and EventBridge monitoring (AWS/MediaConvert namespace with Queue ARN dimension, JobsErroredCount alarm, userMetadata correlation, callbackUrl webhook pattern) — each contains sharp edges that cause silent job failures, duplicate outputs, or unexpected costs. The IAM role attached to a job must have a trust policy listing mediaconvert.amazonaws.com as a principal — without it, job creation succeeds but the job goes immediately to ERROR with INVALID_INPUT_FILE_ACCESS, which looks like an S3 permissions problem but is actually a role assumption failure. S3 event notifications are at-least-once, which means the same video upload can trigger the job-creation Lambda twice — without a DynamoDB conditional write keyed on the S3 object ETag, you pay for duplicate transcoding and produce duplicate output files that overwrite each other. HLS adaptive bitrate switching requires all renditions to have identical segment boundaries — setting GopClosedCadence: 1 in every rendition's H264Settings forces IDR frames at each segment boundary; without it, players stutter or fail to switch quality levels mid-playback. MediaConvert error code 1040 means corrupt or unsupported input format (do not retry), while 3450 and 3451 mean S3 access denied on input and output respectively (fix IAM, then retry). This guide synthesizes all five topics into three structural patterns: job setup and template infrastructure, the event-driven S3 Lambda workflow with idempotency, and HLS output with production monitoring.
TL;DR
- Job setup and IAM: retrieve your account-specific endpoint once with
aws mediaconvert describe-endpoints --region us-east-1 --query 'Endpoints[0].Url' --output textand cache it — all subsequent calls require--endpoint-urlwith this URL; calling the generic endpoint returns HTTP 404. Create an IAM role with trust policy listingmediaconvert.amazonaws.comas principal and inline policy grantings3:GetObjecton your input bucket ands3:PutObjecton your output bucket — wrong trust policy produces INVALID_INPUT_FILE_ACCESS immediately after job submission. Create job templates to avoid reconstructing 200-line settings JSON per job: create custom presets withcreate-preset(codec, bitrate, resolution, container for a single rendition), then create a job template withcreate-job-templatereferencing those presets across output groups. Submit jobs with--job-template ARNand provide only theInputsplus one output group with the per-jobDestination— merge is positional, so the first group in your settings overrides the first group in the template. Browselist-presets --list-by SYSTEMbefore building custom presets; AWS maintains a library covering common web HD/SD and mobile formats. Use on-demand queues (pay-per-minute) for burst workloads and reserved queues (fixed hourly rate) only when processing more than ~6 hours of video per day. Never callget-jobin a polling loop — use EventBridge for completion detection. - S3 trigger workflow and idempotency: configure S3
ObjectCreated:*event notifications on your input prefix filtering by video suffixes (add separateLambdaFunctionConfigurationsentries for each suffix since S3 supports only one suffix per filter rule, or use EventBridge S3 notifications for multi-suffix matching). In the job-creation Lambda, cache the MediaConvert endpoint URL in module scope so it survives warm starts. Before callingcreate-job, write a DynamoDB record keyed on the S3 ETag with aConditionExpression: attribute_not_exists(etag)guard — if the insert fails withConditionalCheckFailedException, this ETag was already processed; skip without creating a duplicate job. Lambda execution role needsmediaconvert:CreateJob,mediaconvert:DescribeEndpoints,iam:PassRolescoped to the MediaConvert role ARN, anddynamodb:PutItem— missingiam:PassRoleproducesAccessDeniedExceptionat job creation despite correct MediaConvert permissions. Derive the output destination path from the input key (e.g.,uploads/user-123/video.mp4→processed/user-123/video/) so outputs land predictably. Handle errors by subscribing an EventBridge rule todetail.status: ERROR— checkerrorCode: 3450/3451 mean S3 access denied (transient, retry after fixing IAM); 1040 means invalid input (permanent, notify user and do not retry). IncludeuserMetadatain everycreate-jobcall withuserId,uploadId, and optionally acallbackUrl— this metadata flows through to every EventBridge event so the completion Lambda can update the database and notify the user without a secondary DynamoDB lookup. - HLS output and production monitoring: use
HLS_GROUP_SETTINGSoutput group withSegmentLength: 6,MinSegmentLength: 0,DirectoryStructure: SINGLE_DIRECTORY, andOutputSelection: MANIFESTS_AND_SEGMENTS. Add one output per rendition — a practical ladder for web delivery is 1080p at 5 Mbps QVBR 8, 720p at 3 Mbps QVBR 7, 480p at 1.5 Mbps QVBR 7. Critical: setGopClosedCadence: 1andGopSize: 180(frames at 30fps) in every rendition'sH264Settings— GopSize in frames is stable across input frame rate variations; setting it in seconds can produce non-integer frame counts at GOP boundaries causing segment misalignment and player stutter on quality switches. Add a second output group of typeFILE_GROUP_SETTINGSwith aFRAME_CAPTUREcodec output (FramerateNumerator: 1,FramerateDenominator: 30,MaxCaptures: 3) to extract thumbnails from the same job. Serve HLS via CloudFront with two cache behaviors:*.m3u8manifests at 30s TTL (static for VOD but players expect freshness) and*.tssegments at 86400s TTL (immutable once written). Create a CloudWatch alarm onAWS/MediaConvertnamespace metricJobsErroredCountwith dimensionQueue: <queue-ARN>, threshold 1, period 300s,TreatMissingData: notBreaching— silent queues with no jobs should not trigger alerts. Add a second alarm onTranscodedBillableMinutesto catch runaway jobs processing unexpectedly long inputs before the billing cycle closes. Validate input duration before callingcreate-jobif users can upload arbitrary-length video — a 4-hour raw recording transcoding to multiple renditions is a large unexpected charge.
Pattern 1 — Job Setup and IAM Configuration
The account-specific endpoint requirement
MediaConvert's most counterintuitive property is the per-account API endpoint. Unlike every other AWS service where you call a regional hostname (e.g., s3.us-east-1.amazonaws.com), MediaConvert assigns each account a unique endpoint that looks like https://abc123def456.mediaconvert.us-east-1.amazonaws.com. Every API call — job creation, queue management, preset creation, job status — must target this account-specific URL. Calling the generic mediaconvert.us-east-1.amazonaws.com returns HTTP 404 Not Found with no error message explaining the issue — it looks like a network connectivity or DNS problem, not an API routing problem.
# Retrieve your account-specific endpoint once
aws mediaconvert describe-endpoints \
--region us-east-1 \
--query 'Endpoints[0].Url' \
--output text
# Returns: https://abc123def456.mediaconvert.us-east-1.amazonaws.com
# Cache in Lambda module scope (survives warm starts)
import boto3, os
_MC_ENDPOINT = None
def get_mc_endpoint():
global _MC_ENDPOINT
if _MC_ENDPOINT is None:
client = boto3.client("mediaconvert", region_name=os.environ["AWS_REGION"])
_MC_ENDPOINT = client.describe_endpoints()["Endpoints"][0]["Url"]
return _MC_ENDPOINT
# All subsequent SDK calls use endpoint_url:
mc = boto3.client("mediaconvert", region_name=region, endpoint_url=get_mc_endpoint())
The endpoint URL does not change between calls for an account — fetch it once at Lambda cold start and reuse it for all subsequent invocations in that execution environment. In the AWS SDK for Python, pass endpoint_url to the MediaConvert client constructor. In the JavaScript SDK v3, pass endpoint in the client config. In CloudFormation or CDK custom resources, call describe-endpoints in the custom resource Lambda and pass the URL to downstream resources as a parameter — do not hardcode the endpoint URL in IaC since it varies between accounts and regions.
IAM role for MediaConvert jobs
Every MediaConvert job requires an IAM role that the service assumes during job execution to read from your input S3 bucket and write to your output bucket. This role is attached per-job via the Role parameter in create-job. The role needs two components: a trust policy allowing mediaconvert.amazonaws.com to assume it, and a permissions policy with S3 read on the input bucket and S3 write on the output bucket. The most common mistake is creating a role with correct permissions but missing the trust policy — without it, job creation succeeds with HTTP 200 and returns a job ID, but the job transitions immediately from SUBMITTED to ERROR with the message INVALID_INPUT_FILE_ACCESS, which looks like an S3 permissions problem but is actually a role-assumption failure.
# Trust policy — mediaconvert.amazonaws.com must be listed as principal
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": { "Service": "mediaconvert.amazonaws.com" },
"Action": "sts:AssumeRole"
}]
}
# Permissions policy
{
"Version": "2012-10-17",
"Statement": [
{ "Effect": "Allow", "Action": ["s3:GetObject", "s3:GetObjectAcl"],
"Resource": "arn:aws:s3:::my-video-input/*" },
{ "Effect": "Allow", "Action": ["s3:PutObject", "s3:PutObjectAcl"],
"Resource": "arn:aws:s3:::my-video-output/*" },
{ "Effect": "Allow", "Action": ["s3:ListBucket"],
"Resource": ["arn:aws:s3:::my-video-input", "arn:aws:s3:::my-video-output"] }
]
}
# If buckets use KMS-managed encryption, also add:
# kms:Decrypt on input key ARN
# kms:GenerateDataKey on output key ARN
If your output bucket uses Block Public Access and object ownership enforced-owner mode (the recommended default for new buckets), you do not need s3:PutObjectAcl in the permissions policy — all uploaded objects are owned by the bucket account regardless. For cross-account scenarios where the MediaConvert role is in account A and the output bucket is in account B, the bucket in account B must have a bucket policy granting s3:PutObject to the role ARN in account A — IAM in account A is insufficient for cross-account S3 writes.
Job templates and output presets for reusable pipelines
A minimal MediaConvert job settings JSON is 50–80 lines for a single H.264 MP4 output. For multiple renditions (1080p + 720p + thumbnail), the settings grow to 200+ lines. Job templates eliminate this repetition: you create the output configuration once as a template, and per-job calls supply only the input and destination override. The template + preset hierarchy is: a preset captures one output's codec and container settings; a job template captures one or more output groups, each containing outputs that reference presets. Browse the AWS-managed system preset library first — presets like System-Generic_Hd_Mp4_Avc_Aac_16x9_1920x1080p_24Hz_6Mbps cover common web delivery formats and save the codec parameter research.
# Create a custom 1080p QVBR preset for web delivery
aws mediaconvert create-preset \
--endpoint-url "$ENDPOINT" \
--name "web-1080p-h264-qvbr" \
--settings '{
"ContainerSettings": {
"Container": "MP4",
"Mp4Settings": { "MoovPlacement": "PROGRESSIVE_DOWNLOAD" }
},
"VideoDescription": {
"Width": 1920, "Height": 1080,
"CodecSettings": {
"Codec": "H_264",
"H264Settings": {
"RateControlMode": "QVBR",
"QvbrSettings": { "QvbrQualityLevel": 8 },
"MaxBitrate": 8000000,
"CodecProfile": "HIGH", "CodecLevel": "AUTO",
"FramerateControl": "SPECIFIED",
"FramerateNumerator": 30000, "FramerateDenominator": 1001,
"SceneChangeDetect": "ENABLED", "AdaptiveQuantization": "HIGH"
}
}
},
"AudioDescriptions": [{
"AudioSourceName": "Audio Selector 1",
"CodecSettings": {
"Codec": "AAC",
"AacSettings": { "Bitrate": 128000, "CodingMode": "CODING_MODE_2_0", "SampleRate": 48000 }
}
}]
}'
# Create a job template referencing two presets
aws mediaconvert create-job-template \
--endpoint-url "$ENDPOINT" \
--name "web-delivery-package" \
--settings '{
"OutputGroups": [{
"OutputGroupSettings": {
"Type": "FILE_GROUP_SETTINGS",
"FileGroupSettings": { "Destination": "s3://PLACEHOLDER/" }
},
"Outputs": [
{ "Preset": "web-1080p-h264-qvbr", "NameModifier": "_1080p" },
{ "Preset": "web-720p-h264-qvbr", "NameModifier": "_720p" }
]
}]
}'
# Per-job call: only supply Input and destination override
aws mediaconvert create-job \
--endpoint-url "$ENDPOINT" \
--role "$ROLE_ARN" \
--job-template "arn:aws:mediaconvert:us-east-1:123456789012:jobTemplates/web-delivery-package" \
--settings '{
"Inputs": [{
"FileInput": "s3://my-video-input/uploads/video.mp4",
"AudioSelectors": { "Audio Selector 1": { "DefaultSelection": "DEFAULT" } },
"VideoSelector": {}
}],
"OutputGroups": [{
"OutputGroupSettings": {
"FileGroupSettings": { "Destination": "s3://my-video-output/processed/video-001/" }
}
}]
}'
Template output group merging is positional — the first group in your create-job settings overrides the first group in the template. This is how per-job destination overrides work: only include the Destination field you want to change; all codec settings from the template carry forward unchanged. A common mistake is omitting the OutputGroups array from the create-job call entirely — without it, the destination override does not apply and all outputs go to the template's placeholder path. Use update-job-template to modify existing templates; changes apply only to new jobs — in-flight jobs at submission time use a frozen snapshot of the template settings. For rollback capability, use versioned template names (web-delivery-v2) and update your code to reference the new name, keeping the old template for retrying in-flight jobs. Use QVBR rate control (Quality-Defined Variable Bitrate) over CBR for web delivery: QVBR adjusts bitrate per-scene, producing smaller files at equivalent perceptual quality — a 90-minute film at QVBR 8 / 8 Mbps max typically averages 3–5 Mbps versus a fixed 6 Mbps CBR stream. Set MoovPlacement: PROGRESSIVE_DOWNLOAD in MP4 container settings to move the moov atom to the beginning of the file, enabling HTTP range-request seeking before the full download completes.
Pattern 1 failure modes
| Failure | Symptom | Fix |
|---|---|---|
| 404 on every API call | create-job, list-queues, list-presets all return HTTP 404 | Calling the generic regional endpoint — run describe-endpoints and pass the result URL via --endpoint-url for all subsequent calls; the generic endpoint does not route to your account |
| Job SUBMITTED → ERROR immediately | Job reaches ERROR within seconds with INVALID_INPUT_FILE_ACCESS | IAM role trust policy is missing or does not list mediaconvert.amazonaws.com as principal — MediaConvert cannot assume the role to access S3; fix the trust policy, not the permissions policy |
| Destination override not applied | Output files land in template's placeholder S3 path not per-job path | Template group merge is positional — include at least one OutputGroups entry with only the FileGroupSettings.Destination field in the create-job settings; omitting the array entirely skips the override |
| Preset not found at job creation | ResourceNotFoundException referencing preset name | Preset names are case-sensitive and region-scoped — verify the preset exists in the same AWS region as the job; custom presets are not replicated across regions automatically |
| Output MP4 has no audio track | Video plays silently despite audio present in input | Each output's AudioDescriptions array must be non-empty — presets created without explicit audio settings produce video-only outputs; add an AAC AudioDescriptions entry referencing Audio Selector 1 |
Pattern 2 — S3 Trigger Workflow and Idempotency
S3 event notification to Lambda trigger
The canonical MediaConvert workflow is event-driven: raw video lands in an S3 input bucket → S3 event notification triggers a Lambda function → Lambda creates a MediaConvert job → EventBridge fires a job state change event when the job completes → a completion Lambda updates your database and notifies users. This pattern costs nothing when idle and scales automatically with upload volume. The first configuration detail that trips teams: S3 event notification suffix filters support only a single suffix per filter rule. If you want to trigger on both .mp4 and .mov uploads, add separate LambdaFunctionConfigurations entries for each suffix pointing to the same Lambda ARN. Alternatively, enable EventBridge integration on the S3 bucket and write a single EventBridge rule with suffix-based content filtering — EventBridge supports matching multiple suffixes in one rule via the suffix comparison operator on the detail.object.key field.
# Lambda resource policy — allow S3 to invoke the function
aws lambda add-permission \
--function-name mediaconvert-job-creator \
--statement-id s3-trigger \
--action lambda:InvokeFunction \
--principal s3.amazonaws.com \
--source-arn arn:aws:s3:::my-video-input \
--source-account 123456789012
# S3 bucket notification for .mp4 uploads
aws s3api put-bucket-notification-configuration \
--bucket my-video-input \
--notification-configuration '{
"LambdaFunctionConfigurations": [
{
"Id": "mp4-trigger",
"LambdaFunctionArn": "arn:aws:lambda:us-east-1:123456789012:function:mediaconvert-job-creator",
"Events": ["s3:ObjectCreated:*"],
"Filter": { "Key": { "FilterRules": [
{ "Name": "prefix", "Value": "uploads/" },
{ "Name": "suffix", "Value": ".mp4" }
]}}
},
{
"Id": "mov-trigger",
"LambdaFunctionArn": "arn:aws:lambda:us-east-1:123456789012:function:mediaconvert-job-creator",
"Events": ["s3:ObjectCreated:*"],
"Filter": { "Key": { "FilterRules": [
{ "Name": "prefix", "Value": "uploads/" },
{ "Name": "suffix", "Value": ".mov" }
]}}
}
]
}'
Idempotency guard with DynamoDB ETag conditional write
S3 event notifications are at-least-once delivery. A single upload can trigger the job-creation Lambda twice — for example, if the Lambda execution fails after creating the job but before acknowledging the event, or if S3 delivers the notification twice due to an internal retry. Without idempotency, you get two MediaConvert jobs processing the same input, producing duplicate output files that overwrite each other and generating double the transcoding charges. The idempotency guard uses a DynamoDB conditional write keyed on the S3 object ETag: before calling create-job, attempt to insert a record with ConditionExpression: attribute_not_exists(etag). If the insert fails with ConditionalCheckFailedException, this ETag was already processed in a previous invocation — return without creating a job. The ETag is stable for a given object version and available in the S3 event record, making it a reliable idempotency key.
import boto3, os, json
_MC_ENDPOINT = None
def get_endpoint(region):
global _MC_ENDPOINT
if _MC_ENDPOINT is None:
client = boto3.client("mediaconvert", region_name=region)
_MC_ENDPOINT = client.describe_endpoints()["Endpoints"][0]["Url"]
return _MC_ENDPOINT
def lambda_handler(event, context):
region = os.environ["AWS_REGION"]
role_arn = os.environ["MEDIACONVERT_ROLE_ARN"]
output_bucket = os.environ["OUTPUT_BUCKET"]
ddb = boto3.client("dynamodb")
for record in event["Records"]:
bucket = record["s3"]["bucket"]["name"]
key = record["s3"]["object"]["key"]
etag = record["s3"]["object"]["eTag"].strip('"')
# Idempotency guard — skip if already processed
try:
ddb.put_item(
TableName=os.environ["JOBS_TABLE"],
Item={ "etag": {"S": etag}, "input_key": {"S": key}, "status": {"S": "SUBMITTED"} },
ConditionExpression="attribute_not_exists(etag)"
)
except ddb.exceptions.ConditionalCheckFailedException:
print(f"Already processed {etag}, skipping")
continue
# Derive output prefix from input key
# uploads/user-123/video.mp4 -> processed/user-123/video/
parts = key.split("/")
user_prefix = parts[1] if len(parts) > 2 else "unknown"
input_name = parts[-1].rsplit(".", 1)[0]
output_prefix = f"processed/{user_prefix}/{input_name}/"
endpoint = get_endpoint(region)
mc = boto3.client("mediaconvert", region_name=region, endpoint_url=endpoint)
job = mc.create_job(
Role=role_arn,
UserMetadata={
"inputKey": key,
"etag": etag,
"userId": user_prefix,
"callbackUrl": "" # set per-request if callers provide webhooks
},
Settings={
"Inputs": [{
"FileInput": f"s3://{bucket}/{key}",
"AudioSelectors": { "Audio Selector 1": {"DefaultSelection": "DEFAULT"} },
"VideoSelector": {}
}],
"OutputGroups": [{
"OutputGroupSettings": {
"FileGroupSettings": { "Destination": f"s3://{output_bucket}/{output_prefix}" }
}
}]
}
)
job_id = job["Job"]["Id"]
ddb.update_item(
TableName=os.environ["JOBS_TABLE"],
Key={"etag": {"S": etag}},
UpdateExpression="SET job_id = :jid",
ExpressionAttributeValues={":jid": {"S": job_id}}
)
print(f"Created job {job_id} for {key}")
The Lambda execution role requires mediaconvert:CreateJob, mediaconvert:DescribeEndpoints, iam:PassRole scoped to the MediaConvert role ARN, and dynamodb:PutItem plus dynamodb:UpdateItem on the jobs table. The iam:PassRole permission is non-obvious — creating a MediaConvert job passes the MediaConvert IAM role, and the caller must have explicit permission to do so. Without iam:PassRole, Lambda receives AccessDeniedException at job creation even though its MediaConvert permissions are correct. Scope the iam:PassRole resource to the specific MediaConvert role ARN to follow least privilege: "Resource": "arn:aws:iam::ACCOUNT:role/MediaConvertJobRole".
EventBridge completion and error handling
MediaConvert emits state change events to EventBridge automatically for every job lifecycle transition — no configuration required on the MediaConvert side. Create two EventBridge rules: one matching detail.status: COMPLETE to trigger a completion Lambda, and one matching detail.status: ERROR to route failures to an error handler. The completion event's detail.outputGroupDetails array contains the S3 paths of every output file produced — use these directly rather than reconstructing paths from the input key, since MediaConvert appends NameModifier values to the destination prefix in ways that are not always predictable from the input filename. For error events, check detail.errorCode to classify failures: error code 1040 means the input file is corrupt or uses an unsupported codec — do not retry, notify the user. Error codes 3450 and 3451 mean S3 access denied on input and output respectively — fix the IAM role or bucket policy, then retry by looking up the input key from DynamoDB using the failed job ID. Error codes 4000+ are transcoding errors related to codec or container compatibility — inspect errorMessage for specifics.
# EventBridge rule: job completion handler
aws events put-rule \
--name "mediaconvert-complete" \
--event-pattern '{
"source": ["aws.mediaconvert"],
"detail-type": ["MediaConvert Job State Change"],
"detail": { "status": ["COMPLETE"] }
}' --state ENABLED
# EventBridge rule: job error handler
aws events put-rule \
--name "mediaconvert-error" \
--event-pattern '{
"source": ["aws.mediaconvert"],
"detail-type": ["MediaConvert Job State Change"],
"detail": { "status": ["ERROR"] }
}' --state ENABLED
# Completion Lambda — read output paths from event, update database
def handle_completion(event):
detail = event["detail"]
metadata = detail.get("userMetadata", {})
upload_id = metadata.get("etag", "unknown")
callback_url = metadata.get("callbackUrl", "")
output_paths = [
path
for group in detail.get("outputGroupDetails", [])
for output in group.get("outputDetails", [])
for path in output.get("outputFilePaths", [])
]
update_record(upload_id, "COMPLETE", output_paths)
if callback_url:
# Validate callback_url is an allowed domain before calling
requests.post(callback_url, json={"status": "COMPLETE", "outputs": output_paths})
# Error Lambda — classify and optionally retry
def handle_error(event):
detail = event["detail"]
error_code = detail.get("errorCode", 0)
job_id = detail["jobId"]
if error_code in (3450, 3451, 3500): # S3 access — transient
resubmit_job(job_id) # look up input key from DynamoDB
else: # 1040 = corrupt input, 4000+ = codec error
notify_user_of_failure(job_id, error_code)
Always validate the callbackUrl value from userMetadata before making the HTTP call — users supply this field and an unvalidated URL creates an SSRF vector. Check that the URL matches an allowlist of permitted domains before calling it in the completion Lambda. For PROGRESSING events (useful for streaming progress to users via SSE), the event includes detail.jobProgress.jobPercentComplete — subscribe a separate rule to detail.status: PROGRESSING if you need real-time progress updates.
Pattern 2 failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Duplicate jobs for same upload | Two MediaConvert jobs processing the same S3 object, duplicate output files | S3 events are at-least-once — add DynamoDB conditional write keyed on S3 ETag before calling create-job; ConditionalCheckFailedException = already processed, skip without creating another job |
| iam:PassRole denied at job creation | AccessDeniedException when Lambda calls create-job | Lambda execution role is missing iam:PassRole scoped to the MediaConvert role ARN — add "Action": "iam:PassRole", "Resource": "arn:aws:iam::ACCOUNT:role/MediaConvertJobRole" |
| Completion Lambda never fires | Jobs complete in MediaConvert but EventBridge rule target not invoked | EventBridge rule target Lambda requires a resource policy allowing events.amazonaws.com to invoke it — add via aws lambda add-permission --principal events.amazonaws.com; also verify rule is in the same region as MediaConvert jobs |
| S3 notification triggers on output files | Output files trigger the job-creation Lambda, creating infinite loop | Input bucket and output bucket must be different, or the notification prefix filter must not match the output destination prefix — never put MediaConvert outputs into the same prefix that triggers job creation |
| Error code 1040 retried indefinitely | Corrupt input file creates a retry loop that never succeeds | Check errorCode in the error event before retrying — 1040 is a permanent input failure; add a retry-count cap in DynamoDB and skip retry for 1040, notifying the user with the errorMessage instead |
Pattern 3 — HLS Output and Production Monitoring
HLS output group structure and rendition ladder
HLS (HTTP Live Streaming) is the correct output format when users play video in browsers or mobile apps and you want quality to adapt to their available bandwidth. MediaConvert produces a master manifest (.m3u8), per-rendition variant playlists, and .ts segments from any input format in a single job — no separate segmentation step required. The output group type is HLS_GROUP_SETTINGS. Use SegmentLength: 6 seconds as a starting point — longer segments (10s) reduce S3 object count but hurt startup time and bandwidth adaptation speed on mobile networks; shorter segments (2s) improve adaptation but create excessive S3 objects. Set MinSegmentLength: 0 to allow MediaConvert to close the last segment at the natural end of the content. Use SINGLE_DIRECTORY for DirectoryStructure so all manifests and segments land in the same S3 prefix as flat files, simplifying CloudFront distribution configuration.
aws mediaconvert create-job \
--endpoint-url "$ENDPOINT" --role "$ROLE_ARN" \
--settings '{
"Inputs": [{ "FileInput": "s3://input/uploads/video.mp4",
"AudioSelectors": { "Audio Selector 1": { "DefaultSelection": "DEFAULT" } },
"VideoSelector": {} }],
"OutputGroups": [
{
"Name": "Apple HLS",
"OutputGroupSettings": {
"Type": "HLS_GROUP_SETTINGS",
"HlsGroupSettings": {
"SegmentLength": 6, "MinSegmentLength": 0,
"Destination": "s3://output/hls/video-001/",
"DirectoryStructure": "SINGLE_DIRECTORY",
"OutputSelection": "MANIFESTS_AND_SEGMENTS"
}
},
"Outputs": [
{ "NameModifier": "_1080p",
"ContainerSettings": { "Container": "M3U8", "M3u8Settings": {} },
"VideoDescription": { "Width": 1920, "Height": 1080,
"CodecSettings": { "Codec": "H_264", "H264Settings": {
"RateControlMode": "QVBR", "QvbrSettings": { "QvbrQualityLevel": 8 },
"MaxBitrate": 5000000, "GopSize": 180, "GopClosedCadence": 1,
"FramerateControl": "SPECIFIED",
"FramerateNumerator": 30000, "FramerateDenominator": 1001
}}},
"AudioDescriptions": [{ "AudioSourceName": "Audio Selector 1",
"CodecSettings": { "Codec": "AAC",
"AacSettings": { "Bitrate": 128000, "CodingMode": "CODING_MODE_2_0", "SampleRate": 48000 }}}]
},
{ "NameModifier": "_720p",
"ContainerSettings": { "Container": "M3U8", "M3u8Settings": {} },
"VideoDescription": { "Width": 1280, "Height": 720,
"CodecSettings": { "Codec": "H_264", "H264Settings": {
"RateControlMode": "QVBR", "QvbrSettings": { "QvbrQualityLevel": 7 },
"MaxBitrate": 3000000, "GopSize": 180, "GopClosedCadence": 1,
"FramerateControl": "SPECIFIED",
"FramerateNumerator": 30000, "FramerateDenominator": 1001
}}},
"AudioDescriptions": [{ "AudioSourceName": "Audio Selector 1",
"CodecSettings": { "Codec": "AAC",
"AacSettings": { "Bitrate": 96000, "CodingMode": "CODING_MODE_2_0", "SampleRate": 48000 }}}]
},
{ "NameModifier": "_480p",
"ContainerSettings": { "Container": "M3U8", "M3u8Settings": {} },
"VideoDescription": { "Width": 854, "Height": 480,
"CodecSettings": { "Codec": "H_264", "H264Settings": {
"RateControlMode": "QVBR", "QvbrSettings": { "QvbrQualityLevel": 7 },
"MaxBitrate": 1500000, "GopSize": 180, "GopClosedCadence": 1,
"FramerateControl": "SPECIFIED",
"FramerateNumerator": 30000, "FramerateDenominator": 1001
}}},
"AudioDescriptions": [{ "AudioSourceName": "Audio Selector 1",
"CodecSettings": { "Codec": "AAC",
"AacSettings": { "Bitrate": 64000, "CodingMode": "CODING_MODE_2_0", "SampleRate": 48000 }}}]
}
]
},
{
"Name": "Thumbnails",
"OutputGroupSettings": {
"Type": "FILE_GROUP_SETTINGS",
"FileGroupSettings": { "Destination": "s3://output/thumbnails/video-001/" }
},
"Outputs": [{
"NameModifier": "_thumb",
"ContainerSettings": { "Container": "RAW" },
"VideoDescription": { "Width": 1280, "Height": 720,
"CodecSettings": { "Codec": "FRAME_CAPTURE",
"FrameCaptureSettings": {
"FramerateNumerator": 1, "FramerateDenominator": 30,
"MaxCaptures": 3, "Quality": 80
}}}
}]
}
]
}'
GOP alignment for seamless rendition switching
The most common HLS quality in production is players that play but stutter or fail to switch renditions when bandwidth changes. The root cause is almost always GOP (Group of Pictures) misalignment between renditions. For adaptive bitrate switching to work, all renditions must have identical segment boundaries — a player switching from 720p to 1080p mid-playback needs the 1080p stream to start at exactly the same timestamp where it left off in the 720p stream. MediaConvert guarantees this via two settings: GopClosedCadence: 1 forces an IDR (Instantaneous Decoder Refresh) frame at every segment boundary, and GopSize set in frames (not seconds) controls the inter-IDR interval. For 6-second segments at 29.97 fps (30000/1001), GopSize should be 180 frames (6 × 30000/1001 ≈ 179.82, round to 180). Setting GopSize in seconds via GopSizeUnits: SECONDS is convenient but can produce non-integer frame counts at segment boundaries — use frames for precise control. Both settings must be identical across all renditions in the output group.
# Correct GOP alignment for 6s segments at 29.97fps
"H264Settings": {
"GopSize": 180, # frames: 6 * 30000/1001 ≈ 180
"GopClosedCadence": 1, # IDR at every segment boundary
"IdrInterval": 0, # let MediaConvert handle IDR placement
"FramerateControl": "SPECIFIED",
"FramerateNumerator": 30000,
"FramerateDenominator": 1001
}
# Apply identical GopSize and GopClosedCadence to EVERY rendition
# Different GopSize values across renditions = segment misalignment = player stutter
# For thumbnail output (separate FILE_GROUP output group in same job):
"FrameCaptureSettings": {
"FramerateNumerator": 1,
"FramerateDenominator": 30, # 1 frame per 30 seconds of content
"MaxCaptures": 3, # max 3 thumbnails per video
"Quality": 80
}
# Output files: video-001_thumb.0000001.jpg, .0000002.jpg, .0000003.jpg
# For a single representative thumbnail: MaxCaptures=1, FramerateDenominator=seconds-into-video
For input video with mixed frame rates (e.g., 24fps film sections mixed with 30fps bumpers), set FramerateConversionAlgorithm: DUPLICATE_DROP to normalize the output to your target frame rate before segmentation. Or set FramerateControl: INITIALIZE_FROM_SOURCE to let MediaConvert match the source frame rate — but in this case, verify that the source frame rate is consistent throughout the file or you may get varying GOP sizes across the output. The thumbnail output group processes concurrently with the HLS output group from the same input decode, so adding thumbnails to a job does not meaningfully increase transcoding time or cost.
CloudFront delivery with split cache behaviors
HLS output from MediaConvert consists of two types of files with very different caching requirements: manifests (.m3u8 files) must be refreshed by players on a short interval to pick up new segment availability, and segments (.ts files) are immutable once written and can be cached indefinitely. A single CloudFront cache behavior with a uniform TTL satisfies neither requirement well — long TTL on manifests causes VOD players to replay stale content lists; short TTL on segments wastes cache capacity and increases S3 origin requests. The correct configuration is two cache behaviors: a *.m3u8 path pattern with 30-second TTL (fast enough for VOD; for live streams reduce to 5-10 seconds), and a default *.ts behavior with long TTL (86400 seconds or more). Use an Origin Access Control (OAC) policy on the S3 origin so segments are served from CloudFront without the S3 bucket being public.
# CloudFront: create distribution with split cache behaviors for HLS
aws cloudfront create-distribution \
--distribution-config '{
"Origins": { "Quantity": 1, "Items": [{
"Id": "hls-s3",
"DomainName": "my-video-output.s3.us-east-1.amazonaws.com",
"OriginAccessControlId": "OAC_ID",
"S3OriginConfig": { "OriginAccessIdentity": "" }
}]},
"DefaultCacheBehavior": {
"ViewerProtocolPolicy": "redirect-to-https",
"CachePolicyId": "MANAGED_CACHING_OPTIMIZED",
"AllowedMethods": { "Quantity": 2, "Items": ["GET", "HEAD"] },
"Compress": true
},
"CacheBehaviors": { "Quantity": 1, "Items": [{
"PathPattern": "*.m3u8",
"ViewerProtocolPolicy": "redirect-to-https",
"DefaultTTL": 30, "MaxTTL": 300, "MinTTL": 0,
"CachePolicyId": "MANAGED_CACHING_DISABLED"
}]},
"Enabled": true
}'
# S3 bucket policy granting CloudFront OAC access
{
"Statement": [{
"Effect": "Allow",
"Principal": { "Service": "cloudfront.amazonaws.com" },
"Action": "s3:GetObject",
"Resource": "arn:aws:s3:::my-video-output/*",
"Condition": {
"StringEquals": { "AWS:SourceArn": "arn:aws:cloudfront::ACCOUNT:distribution/DIST_ID" }
}
}]
}
CloudWatch alarms and EventBridge userMetadata monitoring
MediaConvert publishes six metrics to the AWS/MediaConvert CloudWatch namespace with the queue ARN as the Queue dimension. The two most actionable metrics for pipeline health are JobsErroredCount (any job failure) and TranscodedBillableMinutes (total output duration billed). A JobsErroredCount alarm with threshold 1 and TreatMissingData: notBreaching fires on any failure in a 5-minute window — appropriate for small pipelines where every error matters. A TranscodedBillableMinutes alarm catches runaway jobs processing unexpectedly long inputs before the billing cycle closes. Note: metrics only appear after at least one job has processed in the queue — a new queue with no activity has no datapoints and will not trigger ALARM or OK states until the first job runs. Set TreatMissingData: notBreaching on all alarms to avoid false alerts from idle queues.
QUEUE_ARN="arn:aws:mediaconvert:us-east-1:123456789012:queues/Default"
ALERT_TOPIC="arn:aws:sns:us-east-1:123456789012:ops-alerts"
# Alarm: any job error in 5 minutes
aws cloudwatch put-metric-alarm \
--alarm-name "mediaconvert-errors" \
--namespace "AWS/MediaConvert" \
--metric-name "JobsErroredCount" \
--dimensions "Name=Queue,Value=$QUEUE_ARN" \
--statistic Sum --period 300 --evaluation-periods 1 \
--threshold 1 --comparison-operator GreaterThanOrEqualToThreshold \
--treat-missing-data notBreaching \
--alarm-actions "$ALERT_TOPIC"
# Alarm: billable minutes exceeding expected threshold (catches runaway long inputs)
aws cloudwatch put-metric-alarm \
--alarm-name "mediaconvert-cost-spike" \
--namespace "AWS/MediaConvert" \
--metric-name "TranscodedBillableMinutes" \
--dimensions "Name=Queue,Value=$QUEUE_ARN" \
--statistic Sum --period 300 --evaluation-periods 1 \
--threshold 60 --comparison-operator GreaterThanThreshold \
--treat-missing-data notBreaching \
--alarm-actions "$ALERT_TOPIC"
# EventBridge event structure for correlation via userMetadata:
# {
# "detail": {
# "status": "COMPLETE",
# "jobId": "1696300000000-abc123",
# "userMetadata": {
# "userId": "user-456", "uploadId": "upload-789",
# "callbackUrl": "https://api.example.com/webhook/video-done"
# },
# "outputGroupDetails": [{
# "outputDetails": [{ "outputFilePaths": ["s3://.../_1080p.mp4"] }]
# }]
# }
# }
# userMetadata flows unchanged from create-job to every EventBridge event
# Use it to update the correct DB record without a secondary DynamoDB lookup
# Each key-value pair is a string (max 256 chars) — nested objects are not supported
The userMetadata object — up to 10 string key-value pairs attached at job creation — is the primary correlation mechanism between MediaConvert jobs and your application's data model. Include userId, uploadId, and a callbackUrl when jobs are created on behalf of external callers. The completion Lambda reads these from the EventBridge event detail and can update the database and call the webhook without any additional lookup. MediaConvert does not write job details to CloudWatch Logs — per-job error messages and output paths come from EventBridge events only. For queue depth monitoring (jobs stuck in SUBMITTED state), MediaConvert does not publish a native queue depth metric — implement a scheduled Lambda that calls list-jobs --status SUBMITTED and puts a custom metric to Custom/MediaConvert with period 300 seconds.
Production checklists
Job setup and IAM checklist
- Run
describe-endpointsonce and cache the result — do not call the generic endpoint - IAM role trust policy lists
mediaconvert.amazonaws.comas principal - IAM role permissions policy grants
s3:GetObjecton input bucket ands3:PutObjecton output bucket - For KMS-encrypted buckets:
kms:Decrypton input key,kms:GenerateDataKeyon output key added to role - Custom presets created with
MoovPlacement: PROGRESSIVE_DOWNLOADfor web-seekable MP4 - Job template tested with positional destination override — output files land at expected S3 path
- Template versions named (v1, v2) for rollback capability; old version kept for in-flight job retries
S3 trigger workflow checklist
- Lambda resource policy allows S3 to invoke it (
lambda:InvokeFunctionfors3.amazonaws.comprincipal) - Lambda execution role has
iam:PassRolescoped to the MediaConvert role ARN - DynamoDB jobs table exists with
etagas partition key for idempotency guard - ETag conditional write in place before any
create-jobcall - Input and output buckets are different, or notification prefix filter excludes output prefix
- EventBridge rules for COMPLETE and ERROR added with Lambda resource policies allowing
events.amazonaws.com - Error handler classifies by error code: 1040 = no retry, 3450/3451 = fix IAM then retry
HLS output checklist
GopClosedCadence: 1set on every rendition in the HLS output groupGopSizeexpressed in frames (not seconds), identical value across all renditionsGopSize=SegmentLength × FramerateNumerator / FramerateDenominatorrounded to nearest integer- Each output has a unique
NameModifier— duplicates cause one variant playlist to overwrite another OutputSelection: MANIFESTS_AND_SEGMENTS(notSEGMENTS_ONLY)- Thumbnail output group in same job with
MaxCaptures: 3andQuality: 80
CloudFront and monitoring checklist
- CloudFront distribution uses OAC (Origin Access Control), not legacy OAI
- S3 bucket policy grants
s3:GetObjectto CloudFront OAC — bucket remains private *.m3u8cache behavior with 30s TTL; default behavior for*.tswith 86400s TTL- CloudWatch alarm on
JobsErroredCountthreshold 1 withTreatMissingData: notBreaching - CloudWatch alarm on
TranscodedBillableMinutesfor cost spike detection - Input duration validated before
create-jobif users upload arbitrary-length files callbackUrlin userMetadata validated against allowlist before POST in completion Lambda
Pattern 3 failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Player stutters on quality switch | Video freezes or reloads when bandwidth changes trigger rendition switch | Segment boundaries not aligned — set GopClosedCadence: 1 and identical GopSize (in frames) across all rendition outputs in the HLS output group |
| Master manifest missing some renditions | Player always starts at lowest quality and never upgrades | Each output must have a unique NameModifier — duplicate modifiers cause MediaConvert to overwrite the variant playlist; also verify OutputSelection: MANIFESTS_AND_SEGMENTS not SEGMENTS_ONLY |
| Segments return 403 from CloudFront | .ts files return HTTP 403 despite existing in S3 | S3 bucket does not have a bucket policy granting CloudFront OAC access — add a bucket policy with Principal: { "Service": "cloudfront.amazonaws.com" } and Condition: AWS:SourceArn pointing to the CloudFront distribution ARN |
| CloudWatch alarm stuck in INSUFFICIENT_DATA | Error rate alarm never transitions to OK or ALARM | No jobs have processed in the queue — metrics only appear after at least one job; also verify the Queue dimension value is the full ARN not just the queue name; set TreatMissingData: notBreaching |
| Player buffers on mobile despite good signal | Video buffers at expected bandwidth | Reduce SegmentLength from 10 to 6 seconds for faster startup and adaptation; also verify MaxBitrate is not more than ~10% above target average to prevent peak-segment bursts causing buffer stalls |
| CloudFront serving stale manifest for VOD | Player replays old segment list despite new content available | Set DefaultTTL: 30 and MaxTTL: 300 on the *.m3u8 cache behavior; also send Cache-Control: max-age=30 headers on manifest objects from S3 |
Consolidated failure modes table
| # | Failure | Symptom | Fix |
|---|---|---|---|
| 1 | Generic endpoint returns 404 | All MediaConvert API calls fail with HTTP 404 Not Found | Run describe-endpoints and pass the account-specific URL via --endpoint-url; generic endpoint does not route to your account |
| 2 | Job SUBMITTED → ERROR immediately | INVALID_INPUT_FILE_ACCESS within seconds of submission | Trust policy missing mediaconvert.amazonaws.com principal — fix trust policy, not permissions |
| 3 | S3 read access denied in job | Job errors with code 3450 mentioning Access Denied on input URI | Role lacks s3:GetObject on input bucket prefix; check for bucket policies that conflict with IAM |
| 4 | Destination override ignored | Outputs go to template's placeholder path | Positional merge requires at least one OutputGroups entry with Destination in create-job settings |
| 5 | Duplicate jobs from same upload | Two jobs processing the same S3 object | DynamoDB ETag conditional write guard before create-job |
| 6 | iam:PassRole denied | AccessDeniedException at job creation | Add iam:PassRole scoped to MediaConvert role ARN to Lambda execution role |
| 7 | Completion Lambda never fires | Jobs complete but EventBridge target not invoked | Lambda resource policy missing events.amazonaws.com permission; verify rule and Lambda in same region |
| 8 | S3 notification loop | Output files trigger ingestion Lambda creating infinite jobs | Use separate input and output buckets or exclude output prefix from notification filter |
| 9 | Error 1040 retried indefinitely | Corrupt input keeps generating new failing jobs | Classify error codes before retrying — 1040 is permanent (corrupt/unsupported input); notify user, do not retry |
| 10 | HLS rendition switching failure | Player stutters or freezes on quality change | GopClosedCadence: 1 and identical frame-count GopSize required across all renditions |
| 11 | Master manifest missing renditions | Player starts at lowest quality, never upgrades | Each output needs unique NameModifier; use OutputSelection: MANIFESTS_AND_SEGMENTS |
| 12 | CloudFront segments 403 | .ts files return HTTP 403 | Add bucket policy granting CloudFront OAC s3:GetObject on output bucket |
| 13 | CloudWatch alarm INSUFFICIENT_DATA | Error alarm never transitions to OK or ALARM | Metrics appear only after first job; set TreatMissingData: notBreaching; verify Queue dimension is full ARN |
| 14 | Audio sync drift in long videos | Audio desynchronizes from video after 10–15 minutes of HLS playback | Use identical SegmentLength for audio and video outputs; set AudioGroupId to group audio with its video rendition |
| 15 | callbackUrl SSRF in completion Lambda | Completion Lambda posts to attacker-controlled URLs via userMetadata | Validate callbackUrl against an allowlist of permitted domains before calling; userMetadata is user-controlled input |