The post about Netflix already spoiled a few requirements for synchronized media accessibility, but today we are looking at and listening to the full picture, in sync! You can also watch the recording on YouTube.
Here’s a little challenge for you: The study groups always start as live broadcasts and are later published as recordings. While you read, think about how the study groups would measure up to which test ID, at these different stages.
This test ID ensures that accurate, time-synchronized captions are provided for all pre-recorded audio content within synchronized media, except when the media serves as a clearly labeled alternative to text. The captions must include all dialogue plus equivalents for non-dialogue audio information, such as sound effects, music cues, and speaker identification, without obscuring relevant visual content on the screen.
Failing examples: Captions are present but lack time synchronization, or fail to identify speakers.
Transcripts aren’t captions: Even if a separate transcript is available, alone it does not satisfy this specific requirement.
Questions to help determine if a sound is relevant (Throwback to 16.A. Audio-Only):
Does the sound add pertinent information or cues? Some sounds contribute to setting the scene or adding additional explanation to the narration or dialogue. If so, the sound is relevant and should be included in the transcript.
If the sound was not part of the transcript, would a user lose some information, or would their experience change? If so, the sound is relevant and should be included in the transcript.
Is the sound irrelevant? While it is not required to include sounds in the transcript that convey no meaning, such as music playing in the background, relevant sounds must be included. An instance where music is relevant and needs to be included in the transcript is when the song “Jailhouse Rock” is heard during an interview with its performer, Elvis Presley.
Audio Descriptions (as we’ve learned from Netflix!) are a narration track that describes important visual details that cannot be understood from the main soundtrack alone. These descriptions must fit into natural pauses in the dialogue or narration to ensure they do not overlap with spoken content. While documentaries with heavy narration may require little additional descriptions, any scene where visual context is lost without narration fails 17.B.
Again: A simple text transcript is not enough; the description must be an integrated part of the synchronized audio experience to pass.
Unlike pre-recorded content, there is some leeway for mistakes. For example, words being replaced with [unintelligible] because the live transcribers couldn’t understand what the speaker was saying, or minor spelling mistakes. The captions must still be generally accurate, synchronized with the speaker, and equivalent enough that they do not significantly impact the viewer’s understanding of the program.
Note why live captions are judged differently compared to pre-recorded media: It is assumed that human transcribers will be doing the captions. Today, that is not always the case. While auto-generated captions still benefit from the grace for minor discrepancies, in my opinion, they should be held to stricter standards to incentivise procurement of software with high accuracy.
Attention: Now Leaving WCAG Territory
The last 4 test IDs are based on Section 508 instead of WCAG success criteria. Below are the ICT Standards and Guidelines we will need.
Section 508 503.4 User Controls for Captions and Audio Description – Where ICT displays video with synchronized audio, ICT shall provide user controls for closed captions and audio descriptions conforming to 503.4.
Section 508 503.4.1 Caption Controls – Where user controls are provided for volume adjustment, ICT shall provide user controls for the selection of captions at the same menu level as the user controls for volume or program selection.
Section 508 503.4.2 Audio Description Controls – Where user controls are provided for program selection, ICT shall provide user controls for the selection of audio descriptions at the same menu level as the user controls for volume or program selection.
Any media player presenting synchronized video with audio must provide user controls that allow viewers to enable or disable closed captions. These controls must be functional and selectable (also called: “buttons must do their job”). This test ID only checks for controls, not the accuracy of captions.
Ironically, even if a video has open captions baked in, the standards still mandate separate user controls for captions within the player interface.
Similar to caption controls, this test ID does the same with user controls for audio descriptions. The control mechanism must be activatable and work as expected to toggle the descriptive narration on and off. If the option is missing entirely, disabled, or non-functional when selected, the media player fails.
Continuing from the test process of 17.D., 17.F. we check where the CC controls are. User controls for captions must appear at the same menu level as other primary controls, like volume adjustment or program selection.
The goal is to prevent accessibility features from being harder to find than basic playback functions. In practice, this means the “CC” button should be just as prominent as the “Volume” button.
17.G. is the same process as 17.F., only this time we apply it to the audio descriptions instead. Again: Audio description controls must be at the same level as audio controls or program selection. It’s okay if audio controls are on the main menu and the program selection and audio description are both in a submenu, but the audio description (and captions) cannot be relegated to a submenu alone.
ANDI Accessibility Testing Tool (bookmarklet)
Read more about the study group below, and join the next session through GDG Vienna.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.