Norva

The Complete Guide to Caption Accessibility

A complete framework for evaluating caption content, synchronisation, readability, customisation, controls, devices, and viewer outcomes.

In short: Accessible captions must be discoverable, selectable, synchronised, readable, and sufficiently informative for the viewer's goal. Evaluate dialogue, speaker changes, meaningful sounds, music, timing, size, contrast, background, placement, focus and remote controls, and behavior across supported contexts. Ask viewers what works; never infer needs from history.

Captions provide a text alternative for audio information. They may share timed-text technology with subtitles, but the intended content can be broader than dialogue translation. Actual coverage depends on the supplied caption resource.

Start with the viewer's outcome

Ask whether the viewer needs all spoken dialogue, speaker identification, relevant sound effects, music information, or a combination. Do not request a diagnosis.

Define the viewing context: phone, browser, TV distance, room lighting, shared screen, and input method. Readability is contextual, so a setting that works on a phone may not work across a room.

Verify that the track is really captions

Read the exact label, then sample:

Do not assume every same-language subtitle track provides caption coverage. The guide on captions versus subtitles explains the functional distinction.

Evaluate content quality without overclaiming

Record whether the sample represents speech accurately enough for the task, identifies speakers when needed, and includes meaningful audio cues. Avoid demanding a cue for every sound; relevance depends on context and editorial judgement.

Sample several scenes before calling a title complete or incomplete. A local audit does not establish the quality of an entire catalogue.

Check synchronisation

Captions should appear in a usable relationship with the audio and remain on screen long enough to read. Record early or late patterns across multiple cues rather than relying on one impression.

Separate constant offset, progressive drift, and isolated cue errors. Keep playback speed, item, version, device, and output context stable.

Choose readable size

Text must be large enough for the viewer's distance and vision without covering essential visual information or producing unnecessarily dense lines. There is no universal pixel value for every display.

Use the readable caption-size guide to test actual scenes and viewer outcomes.

Evaluate contrast and background

Caption text needs sufficient separation from changing video frames. Sample bright, dark, detailed, and moving backgrounds. A background box, outline, or shadow may improve separation, but must not obscure critical content.

Use the caption contrast evaluation and caption background guide as paired tools.

Check placement and occlusion

Captions should avoid covering essential on-screen information where supported placement allows it. Test faces, signs, lower-thirds, interfaces, and action near the usual caption area.

Do not move cues randomly from line to line; stable placement can help track speakers. Evaluate the actual supplied positioning and player controls.

Test controls and state

The viewer should be able to find the caption selector, distinguish the current state, select a track, close the menu, and return to playback using intended inputs. Test pointer, keyboard, touch, and TV remote where applicable.

Focus should remain visible and return predictably. Selection should not rely on color alone. Recheck state after resume, episode, version, profile, device, and eligible offline-context changes without assuming universal persistence.

Original evidence: nine-part scorecard

DimensionRepresentative taskResultEvidence
DiscoveryFind and identify caption trackPass/issueExact label
ContentDialogue, speaker, sound, musicResultTimestamps
TimingBeginning, middle, end cuesResultOffset notes
ReadabilitySize, contrast, backgroundResultFrame samples
PlacementAvoid critical contentResultScene notes
ControlSelect and return with intended inputResultFocus path
PersistenceReopen one relevant boundaryResultPaired state

Record “not tested” rather than guessing.

Separate resource and player issues

Missing dialogue or sound information may belong to the supplied caption resource. Clipping, invisible focus, or a state that cannot be selected may belong to player presentation. Evidence can show the affected layer without assigning a root cause.

Norva's current official features and support remain the source of truth for supported controls and devices.

Report access barriers safely

Include exact track label, item/version, device, steps, timestamps, expected outcome, observed outcome, and representative screenshots. Redact credentials, source addresses, account email, and private history. Do not attach media or full caption files.

Common mistakes and limitations

Avoid testing only dialogue, prescribing one size, checking contrast on one frame, relying on mouse input, and treating subtitles as guaranteed captions.

The source supplies caption content. A player can present supported tracks and controls but cannot manufacture missing editorial information.

Frequently asked questions

Are captions only for people who cannot hear audio?

No. Many viewers can benefit, but ask each person what outcome they need rather than assuming.

Is one readable frame enough to prove contrast?

No. Video backgrounds change, so sample bright, dark, detailed, and moving scenes.

Should captions always have a solid background?

Not universally. Choose a background, outline, or other supported treatment based on legibility and visual occlusion in representative scenes.

Can a player fix missing caption content?

It cannot invent cues absent from the supplied resource. Report content and presentation issues separately.

Your next step

Explore Norva's player features

Sources

Explore Norva's Player Features

Sources