What is the Common Intermediate Format?

Common Intermediate Format (CIF) is a picture format defined for interoperable videoconferencing. It uses 352 × 288 luma samples, 4:2:0 chroma sampling, and supports a maximum picture rate of approximately 29.97 frames per second.

Video + audio tracks
Playable derivative
Video processing decodes timed tracks, transforms them, and encodes a deliverable for a target player. This diagram shows video broadly, not specifically the Common Intermediate Format.

How the Common Intermediate Format works

CIF emerged from early video-conferencing standardization as a compromise that equipment built around different television systems could exchange. Its frame is 352 by 288 luminance samples, paired with 4:2:0 chroma and a maximum picture rate of 30000/1001 frames per second. The spatial dimensions align conveniently with the 625-line family, while the timing follows the roughly 29.97-frame convention. Modern workflows mainly encounter it when decoding, normalizing, or preserving legacy communications material.

Key facts

  1. CIF and QCIF were specified in the December 1990 p × 64 kbit/s revision of H.261. CIF describes picture geometry and sampling rather than a video codec or file container.
  2. The 4:2:0 representation stores chroma at lower horizontal and vertical resolution than luma; converters must use the correct chroma geometry to prevent color shifts.
  3. Square-pixel display of 352 by 288 is about 11:9, not 4:3; legacy systems may attach display-aspect assumptions that must be interpreted rather than inferred from dimensions alone.

When the Common Intermediate Format matters

Expect CIF when importing footage from legacy conferencing, surveillance, or communications equipment. Preserve its intended aspect and frame timing during conversion or the result may appear stretched or uneven.

Common use cases for video

These examples cover video broadly, not specifically the Common Intermediate Format.

  • Preparing uploaded video for web, mobile, connected-TV, social, or editorial playback.
  • Creating clips, thumbnails, captions, alternate aspect ratios, and adaptive renditions.
  • Normalizing camera, screen-recording, and user-generated files into predictable outputs.

Working with video

This guidance covers video broadly, not just the Common Intermediate Format.

A demuxer separates tracks from the container, decoders turn compressed streams into frames or samples, and filters apply spatial or temporal changes. Encoders compress the transformed tracks before a muxer writes the chosen output container.

Video compatibility is the product of codec, container, profile, level, frame rate, color, audio, and subtitles. Validate the complete output on target devices because a playable file on one decoder may fail or look different on another.

What you gain

  • Standardized derivatives make diverse source files playable on target devices.
  • A retained master can feed many resolutions, aspect ratios, codecs, and channels.
  • Automated inspection and transformation make large upload volumes consistent.

What it costs

  • More efficient codecs can lower bitrate at similar quality but usually cost more compute and may have narrower support.
  • Higher resolutions and frame rates preserve more detail and motion while increasing processing and delivery requirements.
  • Fast encoding settings improve throughput but can produce larger files or lower quality than slower analysis.

Before production

  1. Inspect codec, container, dimensions, frame rate, color, audio, and subtitle tracks.
  2. Test visual quality and playback support across the slowest and oldest target devices.
  3. Preserve a suitable master before applying lossy, destructive, or delivery-specific changes.

Turn media knowledge into a working pipeline

Connect uploads, processing, AI, storage, and delivery through one declarative API — with the encoding stack, scaling, and format churn handled for you.

Try Transloadit for free