Captions on one video are clean and well timed while another set is close to unusable. The difference is rarely about effort and mostly about which production process was used.
Three routes produce most captions
Automatic speech recognition generates text directly from audio at near-zero marginal cost. It is fast, always available and makes characteristic errors.
Human captioning, prepared from a transcript or typed live, is far more accurate and costs money and time in proportion to the length of the content.
A hybrid route runs recognition first and has a person correct the output. It is cheaper than full human captioning and better than raw machine text, which is why it is now common.
What machine recognition struggles with
Recognition systems are trained on large collections of recorded speech and perform best on the accents, vocabulary and audio conditions best represented in that training material.
Accuracy falls with overlapping speakers, background noise, technical vocabulary, proper nouns and accents that appear less often in the training data.
Those are exactly the conditions of live meetings, classrooms and conference panels, which is why automatic captions often perform worst where they are most relied upon.
Timing and formatting matter as much as words
A caption that is accurate but arrives late, sits over on-screen text, or appears in blocks too long to read is still difficult to use.
Speaker identification, descriptions of meaningful sound and line breaks placed at natural phrase boundaries all affect comprehension, and none is produced reliably by transcription alone.
This is why professional captioning is treated as an editorial task rather than a typing task, and why published quality standards specify presentation as well as accuracy.
Captions and subtitles are not the same product
Subtitles assume the viewer can hear and translate dialogue. Captions assume the viewer cannot hear and must therefore convey non-speech audio that carries meaning.
The files differ technically as well. Open captions are burned into the video image, while closed captions are separate data the viewer can switch on and often restyle.
Platforms that treat all of this as a single feature tend to ship whichever version is cheapest, which is usually an automatic transcript labeled as captions.
Where obligations come from
Requirements arise from a mix of accessibility law, procurement rules, broadcast regulation and platform policy, and they apply differently to public bodies, schools, broadcasters and private sites.
Technical standards published by accessibility bodies are frequently referenced in those rules and specify what must be provided rather than how it must be produced.
Because obligations vary by jurisdiction, sector and distribution method, organizations working out what applies to them generally need specific legal or accessibility advice rather than a general rule.