What are Captions?

Captions are time-aligned text tracks that convey spoken dialogue along with speaker identity and meaningful non-speech sounds a viewer needs to follow the content. They can be embedded in media or delivered as a separate selectable resource.

File + supplied context
Structured media record
Metadata is read, normalized, and used to drive decisions about the media it describes. This diagram shows metadata broadly, not specifically Captions.

How Captions work

A caption file or stream is organized into cues containing timing, text, and sometimes positioning or style instructions. Unlike a transcript, it participates in the playback timeline; unlike dialogue-only subtitles, captioning conveys speakers, music, and meaningful non-speech audio, which is what qualifies a track for accessibility use. Authoring occurs after transcription or translation, followed by timing review, format conversion, packaging, and player validation across the intended languages and devices.

Key facts

  1. WebVTT is widely used for web text tracks, while TTML-family documents are common in professional exchange and streaming; conversion can lose styling or placement not shared by both models.
  2. Overlapping cues, line wrapping, and safe-area placement must be tested against real video because syntactically valid captions can still obscure important graphics or one another.
  3. Burned-in captions become image pixels and work without track support, but viewers cannot disable them, search them as text, or independently select another caption language.

When Captions matter

Add captions to support accessibility, silent viewing, localization, indexing, and search. Incorrect timing or character encoding can make an otherwise valid track difficult or impossible to follow.

Common use cases for metadata

These examples cover metadata broadly, not specifically Captions.

  • Filtering files by dimensions, duration, codec, MIME type, language, or detected content.
  • Building catalogs with searchable descriptions, rights, locations, and relationships.
  • Driving output paths, transformation parameters, moderation, and retention rules.

Working with metadata

This guidance covers metadata broadly, not just Captions.

A metadata reader parses known structures and can derive additional properties from the encoded content. The workflow then validates and normalizes fields before using them for search, routing, naming, filtering, or access decisions.

Metadata can be embedded in a file, stored beside it, or derived during analysis. Track its source and normalization rules, and decide which fields are authoritative, searchable, privacy-sensitive, or safe to copy into derivatives.

What you gain

  • Structured metadata makes media searchable, filterable, and automatable.
  • Technical properties let workflows choose valid transformations before processing.
  • Provenance and rights fields support governance throughout an asset’s lifecycle.

What it costs

  • Copying all metadata preserves context but can leak private or obsolete information.
  • Derived labels scale classification but carry confidence limits and model bias.
  • Rigid schemas improve consistency while making novel or vendor-specific fields harder to retain.

Before production

  1. Distinguish supplied metadata from values detected or derived during processing.
  2. Normalize units, time zones, encodings, and controlled vocabularies at ingestion.
  3. Remove sensitive fields before exposing files or metadata to another audience.

Turn media knowledge into a working pipeline

Connect uploads, processing, AI, storage, and delivery through one declarative API — with the encoding stack, scaling, and format churn handled for you.

Try Transloadit for free