What is VMAF?

VMAF is a full-reference perceptual metric developed by Netflix with academic collaborators to predict subjective video quality from reference and distorted sequences. It combines multiple image-quality features into a score.

Video + audio tracks
Playable derivative
Video processing decodes timed tracks, transforms them, and encodes a deliverable for a target player. This diagram shows video broadly, not specifically VMAF.

How VMAF works

VMAF evaluates an impaired sequence against a temporally and spatially corresponding source, extracts several perceptual features, and fuses them through a trained model. This makes it more viewing-oriented than a single pixel-error formula, but the result still reflects the selected model, preprocessing, and pooling method. Encoding pipelines use it for compression and scaling experiments, regression tests, and ladder analysis, then confirm important decisions with scene inspection and playback evidence. Results should not be generalized to packet loss, corruption, or other artifact classes absent from a model’s training data.

Key facts

  1. Reference and distorted frames must represent the same moments and comparable image geometry; offsets, dropped frames, or inconsistent scaling can measure misalignment instead of degradation.
  2. A VMAF result is tied to a particular model and tool configuration. Scores produced with different models or preprocessing settings should not be assumed directly comparable.
  3. A single pooled score can conceal a short, severely damaged scene. Per-frame traces and lower-tail or scene-level summaries help expose localized quality failures.

When VMAF matters

Encoding teams calculate VMAF when comparing codecs, scaling choices, encoder settings, or bitrate ladders. Results depend on the selected model, reference, and artifact types represented in that model’s training data, so one score should not replace playback testing.

Common use cases for video

These examples cover video broadly, not specifically VMAF.

  • Preparing uploaded video for web, mobile, connected-TV, social, or editorial playback.
  • Creating clips, thumbnails, captions, alternate aspect ratios, and adaptive renditions.
  • Normalizing camera, screen-recording, and user-generated files into predictable outputs.

Working with video

This guidance covers video broadly, not just VMAF.

A demuxer separates tracks from the container, decoders turn compressed streams into frames or samples, and filters apply spatial or temporal changes. Encoders compress the transformed tracks before a muxer writes the chosen output container.

Video compatibility is the product of codec, container, profile, level, frame rate, color, audio, and subtitles. Validate the complete output on target devices because a playable file on one decoder may fail or look different on another.

What you gain

  • Standardized derivatives make diverse source files playable on target devices.
  • A retained master can feed many resolutions, aspect ratios, codecs, and channels.
  • Automated inspection and transformation make large upload volumes consistent.

What it costs

  • More efficient codecs can lower bitrate at similar quality but usually cost more compute and may have narrower support.
  • Higher resolutions and frame rates preserve more detail and motion while increasing processing and delivery requirements.
  • Fast encoding settings improve throughput but can produce larger files or lower quality than slower analysis.

Before production

  1. Inspect codec, container, dimensions, frame rate, color, audio, and subtitle tracks.
  2. Test visual quality and playback support across the slowest and oldest target devices.
  3. Preserve a suitable master before applying lossy, destructive, or delivery-specific changes.

Turn media knowledge into a working pipeline

Connect uploads, processing, AI, storage, and delivery through one declarative API — with the encoding stack, scaling, and format churn handled for you.

Try Transloadit for free