Provide captions for video content

SkillFiles & storage

This is a skill for AI agents that handles captions for video content on web pages. It covers video claude skill work such as adding .vtt caption files via <track> elements to prerecorded videos and checking that captions are enabled on third-party embeds like YouTube or Vimeo. It applies to <video> elements, iframe embeds where the page owner controls the content, and audio-only content.

Available today. Use it from your connected AI after setup.

Have an AI agent that can load skills.

Then ask your AI: use the Provide captions for video content skill

What your AI can do with it

  • Add .vtt caption files to prerecorded videos via <track>
  • Apply caption rules to all <video> elements on a page
  • Check that captions are enabled on YouTube and Vimeo iframe embeds
  • Handle third-party video embeds where the page owner controls the content
  • Address audio-only content

Getting started

  1. Have an AI agent that can load skills.
  2. Add the video-captions skill to the agent's available skills.
  3. Point the agent at the web page or code containing video elements or embeds.
  4. Ask the agent to add captions or verify caption setup on the videos.

What this skill tells your AI

The instructions your AI receives, as published by thedaviddias/front-end-checklist in skills/video-captions/SKILL.md and read by ahel’s review.

Approximately 15% of adults have some degree of hearing loss. Captions are essential for deaf and hard-of-hearing users who cannot access audio content. They also benefit users in sound-sensitive environments (libraries, open offices), users watching without headphones in public, non-native speakers, and users with auditory processing disorders. WCAG SC 1.2.2 is a Level AA requirement — its absence is a legal compliance failure under the ADA, EN 301 549, and similar regulations worldwide.

Quick Reference

  • Prerecorded video with audio: synchronized captions required — WCAG 2.1 SC 1.2.2 (Level AA)
  • Live video with audio: real-time captions required — WCAG 2.1 SC 1.2.4 (Level AA)
  • Use <track kind='captions'> with a .vtt (WebVTT) file for HTML5 <video> elements
  • Captions must include all spoken dialogue, speaker identification, and relevant non-speech audio (music, sound effects)
  • Subtitles and captions are different: captions include non-speech audio; subtitles translate dialogue only

Check

Find all <video> elements and video embeds (<iframe> from YouTube, Vimeo, etc.). For each <video> with audio: check for a <track> child element with kind='captions' and a valid src pointing to a .vtt file. Verify the default attribute is present on at least one track so captions are on by default (or document the UX reason they are off by default). For YouTube/Vimeo embeds: check that the platform's caption toggle is accessible. Also check that the .vtt file exists and is valid (not empty, not just music notes).

Fix

For <video> elements without captions: (1) Create a WebVTT (.vtt) file containing synchronized caption text — include all spoken words, speaker IDs for multi-speaker content, and descriptions of relevant sounds (e.g., '[applause]', '[upbeat music]'). (2) Add <track kind='captions' srclang='en' label='English' src='captions-en.vtt' default> inside the <video> element. (3) For auto-generated captions (YouTube, AI tools): review and correct errors — auto-captions average 80% accuracy and often fail on proper nouns, technical terms, and accented speech. (4) For live streams: implement real-time captioning via a third-party captioning service or CART (Communication Access Realtime Translation).

Explain

WCAG 2.1 SC 1.2.2 (Captions — Prerecorded, Level AA) requires synchronized text alternatives for all audio in prerecorded video content. Captions differ from subtitles: captions are intended for deaf/hard-of-hearing viewers and must include non-speech information (sound effects, music), while subtitles translate dialogue for viewers who can hear but do not understand the language. The HTML <track> element with kind='captions' delivers WebVTT files that browsers render as synchronized on-screen text. The kind='subtitles' value is for translation only and does not satisfy SC 1.2.2 because browsers may omit non-speech annotations.

Code Review

Review the rendered markup and interactive states that affect Provide captions for video content. Flag exact elements, roles, labels, focus behavior, or keyboard interactions that violate the rule, and note how to verify the fix with browser accessibility tooling or assistive tech.


For full implementation details, code examples, and framework-specific guidance, see references/rule.md.

Rule page: https://frontendchecklist.io/en/rules/accessibility/video-captions

Signals

GitHub stars
74k
Forks
7k
Last commit
Aug 2026

Questions

When should this skill be used?
It applies to all <video> elements and third-party video embeds (YouTube, Vimeo) where the page owner controls the content, including prerecorded videos needing .vtt caption files and audio-only content.
How are captions added to prerecorded videos?
Prerecorded videos require .vtt caption files, which are attached through a <track> element.
What about videos embedded via iframe?
For videos embedded via <iframe>, check that the video platform captions are enabled.
Does it work with third-party platforms?
Yes, it covers third-party video embeds such as YouTube and Vimeo where the page owner controls the content.
Advanced
Item type
skill
Key
video-captions
Source
github.com/thedaviddias/front-end-checklist